← Back to events
ActiveAI

BenchMIRT: What are LLM benchmarks actually measuring?

Photo: Allen Institute for AI

What happened

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.

Summary assembled by rule from the sources below

Why it's spreading

Timeline

  1. First appeared on Allen Institute for AIAllen Institute for AI
  2. Hugging Face published an announcementHugging Face

Sources