← 返回事件
持续讨论AI

My LLM eval cried wolf. Here's what I measured

图:Hacker News

发生了什么

A case went from 5/5 to 2/5 with nothing changed. How I measured the noise floor of an LLM-judged eval, what it caught the week after, and where it still can't see.

摘要按规则整理自下方来源原文

为什么在扩散

来源

社区讨论