Alibaba President: AI agents can talk, but can they actually do the work?

What happened
The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.
Summary assembled by rule from the sources below