The Implications of Linguistic Illegibility for LLM Security
发生了什么
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fa…
摘要按规则整理自下方来源原文
为什么在扩散
时间线
- Hacker News 最先出现Hacker News
- Lobsters 出现讨论Lobsters