← 返回事件
持续讨论AI

The Implications of Linguistic Illegibility for LLM Security

发生了什么

LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fa…

摘要按规则整理自下方来源原文

为什么在扩散

时间线

  1. Hacker News 最先出现Hacker News
  2. Lobsters 出现讨论Lobsters

来源