Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

What happened
A benchmark for evaluating AI agents on research workflows across scientific domains
Summary assembled by rule from the sources below

A benchmark for evaluating AI agents on research workflows across scientific domains
Summary assembled by rule from the sources below