BenchMIRT: What are LLM benchmarks actually measuring?

- Allen Institute for AI: 4 events in the last 90 days
- Previous: 5 days earlier · Ai2 and Providence Swedish Cancer Institute partner to advance AI-assisted scientific discovery
What happened
BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.
Summary assembled by rule from the sources below
Why it's spreading
Timeline
- First appeared on Allen Institute for AIAllen Institute for AI
- Hugging Face published an announcementHugging Face