← 返回事件
持续讨论AI

OpenAI has a LOT of work to do if they think Luna can compete with Jev

图:Hacker News

发生了什么

We've routed decisions on LLM confidence since 2023, and calibration is what we most hope OpenAI's new Decisions API gets right. The model underneath, GPT-6 Luna, isn't there yet: on 3,600 reasoning problems its 99% meant 68%, and by five inference steps it was at a coin flip.

摘要按规则整理自下方来源原文

为什么在扩散

来源