ARC-AGI-3 Harness Fail: Why GPT-5.6 Sol Scored Only 7.8% – And the Simple Fix That Tripled Its Score

11 days ago 高效码农

GPT-5.6 Sol Scored 7.8% on ARC-AGI-3. The Problem Was the Harness, Not the Model. When we first pulled up the ARC-AGI-3 leaderboard, our immediate reaction was confusion. GPT-5.6 Sol had already solved the cycle double cover conjecture, a longstanding open problem in mathematics. It beat Pokémon FireRed using a vision-only setup. Yet on ARC-AGI-3, a benchmark built from 2D puzzle games, it managed a mere 7.8%. GPT-5.5’s performance was even more extreme, scoring 0.4%, which essentially means it failed to play the games at all. How does a model capable of beating Pokémon get stuck on simple 2D puzzles? The …