The Implications of Linguistic Illegibility for LLM Security
What happened
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fa…
Summary assembled by rule from the sources below
Why it's spreading
Timeline
- First appeared on Hacker NewsHacker News
- Discussion started on LobstersLobsters