← Back to events
ActiveAI

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Photo: Hacker News

What happened

A benchmark for evaluating AI agents on research workflows across scientific domains

Summary assembled by rule from the sources below

Why it's spreading

Sources