The Implications of Linguistic Illegibility for LLM Security
- View PDF HTML (experimental) Abstract:LLMs are trained to generate natural language.
- However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation.
- We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks.
Unverified
- View PDF HTML (experimental) Abstract:LLMs are trained to generate natural language.
- However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation.
- We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks.
Sources: Arxiv