GPT 5.5 is better than Opus 4.8
I've been saying for a long time that GPT 5.5 is better than Opus 4.7 at their highest intelligence settings.
In fact, GPT 5.5 is even better than the new Opus 4.8!
The “traditional” benchmarks that measure this (SWE-Bench Pro) are really bad and don't match the experience you get as a programmer. They put all the models close together (even the worse ones), gave inconsistent results, were contaminated, and had serious flaws.
For example, they let Opus cheat by looking at the git log (wow)
That's why they've created a new benchmark that seems to better represent these differences and produces much more believable results.
I don't think these benchmark results will surprise anyone who's used these models a lot for real-world programming tasks.
Link below ⬇️