← Back to events
ActiveAI

Agent Evaluation Metric for multi-turn conversations

Photo: AWS Machine Learning

What happened

Multi-turn agents fail in ways single-turn evaluation misses: one early mistake corrupts every later turn. This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality, applied to its first dimension, correctness, to pinpoint the turn that caused a failure and separate it from the turns that inherited it.

Summary assembled by rule from the sources below

Why it's spreading

Sources

Official