← Back to events
ActiveAI

OpenAI has a LOT of work to do if they think Luna can compete with Jev

Photo: Hacker News

What happened

We've routed decisions on LLM confidence since 2023, and calibration is what we most hope OpenAI's new Decisions API gets right. The model underneath, GPT-6 Luna, isn't there yet: on 3,600 reasoning problems its 99% meant 68%, and by five inference steps it was at a coin flip.

Summary assembled by rule from the sources below

Why it's spreading

Sources