Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
What happened
Real-SWE benchmarks frontier AI models on private production codebases licensed from real companies. Eight model and harness configurations, ten tasks, 640 scored rollouts.
Summary assembled by rule from the sources below