All posts

Translated from Spanish · See the original

GPT 5.5 is better than Opus 4.8

I've been saying for a long time that GPT 5.5 is better than Opus 4.7 at their highest intelligence settings.

In fact, GPT 5.5 is even better than the new Opus 4.8!

The “traditional” benchmarks that measure this (SWE-Bench Pro) are really bad and don't match the experience you get as a programmer. They put all the models close together (even the worse ones), gave inconsistent results, were contaminated, and had serious flaws.

For example, they let Opus cheat by looking at the git log (wow)

That's why they've created a new benchmark that seems to better represent these differences and produces much more believable results.

I don't think these benchmark results will surprise anyone who's used these models a lot for real-world programming tasks.

Link below ⬇️

deepswe.datacurve.ai

Views: 1.7KLikes: 7Replies: 3Reposts: 1View on X

Enjoyed this post?

Leave me your email and I'll let you know when I publish something new.

Related posts

It’s sooo easy to tell which apps have been vibecoded. So easy. There are details that scream it right in your face from minute one. Little things… but they immediately give you away: 🎨 The design. It’s not that it’s ugly. It’s that it’s the default design. You can almost guess which model has

Views: 93.6KLikes: 802Replies: 64Reposts: 32View on X

If Fable is already almost unusable even when you’re paying €100/€200 a month, I’m scared to see what happens when the extended limits period ends. Anthropic needs compute now. Well, it’s needed it for months, but now people are realizing just how much you can get done with Codex and even Grok.

Views: 10.5KLikes: 155Replies: 7Reposts: 3View on X