← Back to events
ActiveAI

A new skill helps developers find AI agent risks, fix them, and prove it worked

Photo: Microsoft Source

What happened

Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval to prove whether the fix worked.

Summary assembled by rule from the sources below

Why it's spreading

Sources