A new skill helps developers find AI agent risks, fix them, and prove it worked
- Microsoft: 80 events in the last 90 days
- Previous: earlier the same day · We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together [Read more]
What happened
Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval to prove whether the fix worked.
Summary assembled by rule from the sources below