A new skill helps developers find AI agent risks, fix them, and prove it worked
发生了什么
Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval to prove whether the fix worked.
摘要按规则整理自下方来源原文