AI MODELS◐ DEVELOPINGN · Network◆ Concentrating
OpenAI’s AGI benchmark score came from a harness, not the base model
Sep 6, 2026SOURCE: thenextweb.com
SO WHAT
In the thousand-day window, this is a signal that model capability claims are increasingly mediated by scaffolding—where the inversion shifts from “who has intelligence” to “who controls the evaluation surface.”
ARC Prize says its harness ran GPT-6 Astra twice, but OpenAI’s setup returned 99.9% versus 62.7%. Benchmark authors explicitly do not claim AGI and OpenAI revised five published results.
This is Negative Resistance’s reframed reading of a reported signal. The headline and analysis above are our interpretation through the thousand-day-window lens. The original reporting lives at the source linked above.