Claude Sonnet 5.5: Impressive Demos and Some Fair Questions
The cycle of AI updates is getting faster and faster. It’s only been a few weeks since Opus 5.5 and GPT Astra were released, and now Anthropic has released Claude Sonnet 5.5. So, how good is it, and how does it compare to the other models? Let’s take a look at some demos, comparisons, and the reactions they’re getting.
A 40-versus-40 shooter from Higgsfield
Let’s start with this demo from Higgsfield. They say they used Sonnet 5.5 to create a 40-versus-40 shooter, complete with air strikes. And I must say, it looks really impressive. But there’s some doubt in the comments, which I can understand.
The scale is huge, the models look good, and the textures are surprisingly detailed. That has some people questioning how much of this was actually made by Sonnet itself. Where did the models and textures come from? Were other tools involved? Did it build everything from scratch or put together existing assets? Those are different things, and the finished demo alone doesn’t tell us the whole story.
One commenter points out that 40-versus-40, with air strikes on top, is getting a bit ridiculous. Another says it still uses plenty of human-made models. Based Apple is even more skeptical, calling it engagement bait and questioning whether the game could have been made in the time since the model was released.
I find these reactions interesting because the discussion isn’t just people worshipping AI or completely dismissing it. Some are raising practical questions based on their experience with different parts of the process. Getting a game running is one thing. Creating the models, textures, effects, and gameplay systems is another.
We also have Ruben, who shared an example of a project he says took him a month to build using Astra 6, Opus 5.5, and GLM 5.3. His reaction was basically, “Just Sonnet? No way.” That doesn’t prove the Higgsfield demo is fake, but it does explain why people want to see more of the process.
Sonnet 5.5 vs. GPT 6 Astra vs. GPT 6 Sol
Speaking of the competition, here’s a comparison between Sonnet 5.5, GPT 6 Astra, and GPT 6 Sol.
This isn’t a scientific test, and different models can do better or worse depending on the challenge. But in this example, GPT 6 Sol looks like it barely even tried. The result doesn’t seem to capture what the task was asking for.
Between Sonnet 5.5 and GPT Astra, though, I don’t see a huge difference beyond the visual style. You might prefer one over the other, but from this example alone, I wouldn’t call either one a clear winner.
These comparisons are useful for seeing how models approach the same task, but a single result only tells us so much. How many attempts were needed? Did either model require corrections? Does the result actually work, or does it just look good? Those details matter when you’re choosing something to use in your own projects.
A detailed Jeep, but how long did it take?
We have another test, this time from Max, who pushed back against claims that Sonnet had easily beaten Astra. He asked both models to code a 3D Jeep in a single HTML file using Three.js, with exactly the same prompt and reference image.
According to Max, Sonnet used 6% of his weekly quota. The textures were incredibly detailed, but the result took an hour and 25 minutes. Astra used 7% of his weekly quota and finished in 35 minutes. It had slightly less detail, but still looked solid.
So, in that test, Sonnet produced more detail, while Astra finished much sooner. Which one you prefer depends on how much that extra detail matters to you and how long you’re willing to wait.
Max’s own verdict was that he would still rather use Opus 5.5 on medium or high settings for his projects. That’s useful context, because the model that produces the most impressive screenshot isn’t automatically the one you’ll want to use every day. Speed, usage limits, and how much fixing you need to do afterward all matter.
If you’re experimenting with an idea, waiting over an hour for each version could make the process frustrating. If you’re after a detailed result and don’t mind the wait, you might feel differently.
The ocean demo and what interests me as a 3D artist
Here’s another demo from Matthew Berman. Whatever you think about these models, and whichever one you prefer, the progress on display is remarkable. Just look at this 3D ocean simulation.
The visuals are beautiful, but for me, as a 3D artist, the more interesting part is how these systems can work with existing tools. When a model writes code or helps operate an application, it can leave you with a project you can open, inspect, and continue working on.
Think about what that means for workflows involving Blender, Unity, Unreal Engine, or Photoshop. These are applications that take people years to become comfortable with. A beginner might spend months learning enough to put together even one part of an ambitious project.
Being able to help with the setup, write parts of the code, and handle repetitive steps can make a big difference. You can spend more time deciding what you want to create and less time figuring out every technical step needed to get started.
You still need to judge whether the result is good, understand what needs fixing, and guide the project toward something useful. But having help with the technical work can make it easier to try ideas that you might otherwise put off.
That’s what makes this interesting to me. These workflows can use the same tools artists already work with, making it easier to bring the results into an existing project and keep developing them.
Sonnet 5.5 vs. Opus 5.5
Finally, here’s another example from Filipe, comparing Sonnet 5.5 with Opus 5.5. Again, at least for this type of test, the difference in visual quality looks fairly small. A lot of it comes down to which style you prefer.
This example was created in HTML, so the result is something you can open in a browser and edit through its code. If you want to take over, change the behavior, or build on what the model has produced, you have something you can work with.
That’s the part I keep coming back to. The impressive result is one thing, but being able to continue working on it yourself is just as important. For an artist or developer, the value grows when the output fits into a familiar workflow and gives you room to make your own changes.
Sonnet 5.5 has some impressive demos here, but the comparisons also show why it’s worth looking beyond visual detail. How long did the task take? What assets were supplied? How much human input was involved? And how easy is the result to edit?
Those are the questions I’d want answered before deciding which model to use. In the meantime, let me know which comparison you found most convincing and what you’d like to see tested next.
