← Back to events
ActiveAI

My LLM eval cried wolf. Here's what I measured

Photo: Hacker News

What happened

A case went from 5/5 to 2/5 with nothing changed. How I measured the noise floor of an LLM-judged eval, what it caught the week after, and where it still can't see.

Summary assembled by rule from the sources below

Why it's spreading

Sources