Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

- ARC-AGI-3 近 90 天出现 3 次
- 上一次:同一天稍早 · GPT-6 Astra reaches and surpasses human-level performance on ARC-AGI-3 with a memory harness, but still at 1000x the cost - ARC Prize Foundation
发生了什么
OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet doesn't call this proof of AGI, but he does see the progr…
摘要按规则整理自下方来源原文