OpenAI has a LOT of work to do if they think Luna can compete with Jev

- Lot: 2 events in the last 90 days
- Previous: 34 days earlier · That's a Lot of YAML
What happened
We've routed decisions on LLM confidence since 2023, and calibration is what we most hope OpenAI's new Decisions API gets right. The model underneath, GPT-6 Luna, isn't there yet: on 3,600 reasoning problems its 99% meant 68%, and by five inference steps it was at a coin flip.
Summary assembled by rule from the sources below