298points · 1h ago

Sonnet 5.5

anthropic.com·by D2OQZG8l5BI1S06·1h ago

Discussion 193 comments

bestnew
⌘↵ to post · markdown supported
simonw·1h ago
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Here's how the thinking effort levels compare:

  low
  27 input, 1,623 output, thinking_tokens: 0
  1.6284
  Duration: 10138ms (10s)
  
  medium
  27 input, 1,796 output, thinking_tokens: 0
  1.7914 cents
  Duration: 11266ms (11s)

  high
  27 input, 2,334 output, thinking_tokens: 745
  2.3394 cents
  Duration: 17376ms (17s)

  xhigh
  27 input, 5,730 output, thinking_tokens: 2535
  5.7354 cents
  Duration: 41882ms (41s)

  max (failed to return response)
  27 input, 128,000 output, thinking_tokens: 128000
  $1.28
  Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.
0
parkersweb·2m ago
I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
0
thefourthchime·46m ago
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...

2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...

Up until very recently, all models struggled with this.

All results: https://jonclegg.github.io/pacman-bakeoff/

0
ilamont·27s ago
Thank you for doing this. It is very helpful not just for capabilities but also for costs.
0
judge2020·27m ago
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
0
thefourthchime·22m ago
Yes! The point being that up until yesterday, every model struggled with this, and now they don't.
0
russellbeattie·24m ago
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
0
croemer·1h ago
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
0
gumby271·58m ago
If it was trained on HN, there would be a 60% chance of it just saying "I'm so tired of this request, can we please move on"
0
platinumrad·35m ago
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
0
keeeba·27m ago
Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?
0
TomGarden·56m ago
Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
0
petu·51m ago
That's max output tokens per response limit, separate from context length
0
simonw·41m ago
It's the output token limit, which has been 128,000 for Claude models for quite a while note
0
croemer·21m ago
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
0
simonw·12m ago
I think this is a bug. I've not seen this problem from any of the other frontier models.
0
Insanity·4m ago
Do other models put a hard cap on the output tokens it can generate?
0
heyjstn·49m ago
I think the next models will be benchmaxxing on the Pelican benchmark tbh
0
aimaxxed·44m ago
“Pelicans are solved.”
0
Sol-·1h ago
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

0
maherbeg·1h ago
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.

Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?

0
datadrivenangel·37m ago
Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium, so what's the point of sonnet?
0
egeozcan·32m ago
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.

OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.

Totally different uses.

0
phainopepla2·1h ago
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
0
andrepd·12m ago
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?

The proof of the pudding.

0
Imustaskforhelp·1h ago
> It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.

In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.

There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):

> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.

> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.

> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.

What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)

I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.

I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?

Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.

[0]: https://news.ycombinator.com/item?id=49808422

[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...

0
RGS1811·1h ago
> Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.

For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.

0
chrismustcode·1h ago
Cache read is the same as Opus as well where most agentic workflow cost comes from.

Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?

0
gregwebs·1h ago
> I want to retain some semblance of understanding

How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?

I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?

0
afro88·59m ago
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
0
alansaber·1h ago
When they inevitably drop allocation after post-launch hype dies down.
0
losvedir·1h ago
Useful for API requests, when using AI in the product rather than to build the product.
0
neuronexmachina·50m ago
Most business/enterprise accounts also have to pay API rates.
0
losvedir·22m ago
Exactly. I'm saying that Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token.
0
MisterMunchkin·36m ago
It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

And my job won’t even pay for Claude now because it’s so ruinously expensive.

0
throwa356262·24m ago
Obviously not as "intelligent" but almost 10x cheaper

Mimo 2.6 Pro: 0.04/0.4/0.87

Sonnet 5.5: 0.2/2/10

Opus 5.5: Sonnet prices times 2

What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?

0
edu·13m ago
What model are you using ?
0
system2·5m ago
Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
0
zarmin·1m ago
What harness are you using with those?
0
abejora·1h ago
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card

0
eli·1h ago
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
0
abejora·1h ago
You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

0
joeyhage·44m ago
Claude, is that you?
0
subscribed·25m ago
I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

Anthropic made it that way, and I'd say the lower score is accurate.

0
Leary·1h ago
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
0
radlad·1h ago
I believe you meant to cite the Opus 5.5 System Card which states:

> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...

I cannot find a Sonnet 5.5 system card.

0
abejora·1h ago
0
oh_no·1h ago
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

0
manojlds·1h ago
Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
0
MadameMinty·1h ago
That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

0
manojlds·1h ago
Fallback was usually Opus 4.8
0
a13o·4m ago
This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.
0
wongarsu·1h ago
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5

Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models

0
gozzoo·15m ago
what is the easyest way to use the chinese models and which harness does work with them well?
0
Iolaum·12m ago
OpenCode harness with their subscription would be my recommendation.
0
ttul·1h ago
Daybreak Blue is not bad and the bar to get into OpenAI's program is reasonable.
0
watusername·29m ago
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
0
ghoshbishakh·1h ago
So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
0
pebbly_bread·31m ago
Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
0
mnicky·43m ago
Well, it's performance "surface" (is there a better term for this?) is probably very narrow compared to Fable :)
0
heyjstn·53m ago
doom marketing at its finest
0
johnmlussier·1h ago
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.

This is bollocks. Their safeguards are shit.

0
solenoid0937·1h ago
You should read the actual docs for the CVP. At the very top:

https://support.claude.com/en/articles/14604842-real-time-cy...

> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models

You obviously should not expect the CVP to cover this model either.

0
machomaster·1h ago
He did mention Sonnet...
0
solenoid0937·1h ago
It takes about 2 seconds of critical thinking to realize that if Opus 5.5 isn't covered yet, neither will a model that just launched an hour ago.
0
gowld·1h ago
Does it also take 2 seconds of critical thinking to realize that the models that are covered should be accurately named by the people making the decisions?
0
solenoid0937·50m ago
Sure, the documentation should be up to date but it's obviously not? That doesn't excuse not thinking critically.
0
icedchai·1h ago
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
0
dejw·1h ago
it did exactly what a human would do - "I can't look at this shit"
0
K0balt·1h ago
Yuh—- no. 4.8 can handle this bullocks lol
0
jchw·1h ago
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.

IDA Pro and Ghidra, thankfully, still lack such safeguards...

(No other model I've tried has refused either FWIW.)

0
dom96·46m ago
Funnily enough the Fable safeguards are the worst and testing Sonnet 5.5 didn't trigger them as much as it even did for Opus on my benchmarks[1].

1 - https://bench.killswitch-lang.org

0
film42·1h ago
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
0
skeledrew·14m ago
> Switched to OpenCode + OpenRouter

This is the way.

0
jauntywundrkind·1h ago
It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
0
film42·38m ago
Bingo. And to prove your point, after switching to cheap open models (I think Qwen?) it did indeed find a bug in my WAL implementation.
0
nightpool·1h ago
https://support.claude.com/en/articles/14604842-real-time-cy... says that Cyber Verification Program doesn't apply to Opus 5.5 yet, they hope to roll it out for 5.5 "soon"
0
giancarlostoro·1h ago
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.

Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.

0
AshamedBadger56·1h ago
Yup. As far as I can tell, the Cyber Verification Program does absolutely nothing.
0
wkcheng·1h ago
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.

0
ricardobeat·1h ago
At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks worse at xhigh.
0
wkcheng·1h ago
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
0
oh_no·55m ago
i'd love to see them re-enter that space but given haiku 5 never happened I wouldn't bet on it

i think they see what openai charges for luna and just don't want to try and compete

0
water-drummer·27m ago
They did mention in the Opus 5.5 announcement blogpost that Sonnet and Haiku 5.5 will follow soon.
0
canad3nse·34m ago
But they literally stated that they would release Sonnet 5.5 and Haiku 5.5 after Opus 5.5 was released
0
pdantix·31m ago
they've already said in both the opus 5.5 and sonnet 5.5 blog posts that haiku 5.5 is coming
0
delillos·1h ago
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.

Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.

0
usaar333·1h ago
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.

But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.

0
RussianCow·1h ago
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.

With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.

0
Jcampuzano2·1h ago
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.

Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.

I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.

0
SubiculumCode·1h ago
t/s maybe? IDK, because their token speed comparison was against Sonnet 5.
0
solenoid0937·1h ago
It literally does not?
0
dominotw·59m ago
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
0
heyjstn·1h ago
Have anyone tried a workflow that:

- Fable 5.1 for planning/adversarial reviewer

- Opus 5.5 for well-scoped tasks break down

- Sonnet 5.5 for these well-scoped tasks implementation

I think the blocker might be how efficient the context is compacted and sending around between these agents

0
afro88·57m ago
Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.
0
chrismustcode·1h ago
You might as well use Opus for everything there.

Changing model would be cache busting spiking usage for no good reason when Opus can do it all.

Haiku 5.5 might fit well though depending on pricing.

0
SirMadam·54m ago
Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?
0
manquer·7m ago
[delayed]
0
enraged_camel·45m ago
Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
0
mnicky·40m ago
On the contrary, subagents save context overall, when the task is sufficiently large.

Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).

0
Aboutplants·17m ago
Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.
0
Jcampuzano2·1h ago
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.

Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.

Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.

0
nanook·14m ago
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
0
sajithdilshan·26m ago
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
0
ricericerice·12m ago
my feeling is you're most likely wasting money using Opus for implementation. The plan and task breakdown should be specific enough that Sonnet can implement without you noticing a difference.
0
avree·1h ago
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
0
yapfrog·1h ago
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
0
alansaber·1h ago
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
0
gregwebs·1h ago
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.

From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.

OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.

0
system2·8m ago
Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.
0
tombert·1h ago
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
0
aniceperson·31m ago
I saw that with corporate software too. What works is creating a skill with the task steps, it fades its initial reasoning. (I am not talking about observer safe guards, but the safety RTL).
0
dom96·47m ago
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.

https://bench.killswitch-lang.org/

    Claude Sonnet 5    17.8%
    Claude Sonnet 5.5  7.4%
0
onlyrealcuzzo·1h ago
> In our testing, it costs up to 30% less per task than its predecessor.

> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.

They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.

Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.

I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).

For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.

This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

Hopefully they release a Haiku that actually has a reason for existing.

0
jchw·1h ago
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.

DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.

0
velcrovan·1h ago
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
0
eli·1h ago
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
0
jchw·1h ago
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.

MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.

I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.

I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.

Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.

0
eli·1h ago
I read they identified a training bug and were going to push out an updated release to fix the looping. I really like it overall.

GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.

0
cbg0·1h ago
> This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

It's been out for an hour and you've already concluded this?

0
criemen·1h ago
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
0
enraged_camel·1h ago
From TFA:

>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

0
aniceperson·36m ago
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.

My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.

Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?

https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...

0
criemen·16m ago
> but openai went kamikaze

I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.

0
aniceperson·3m ago
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.

I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.

0
mroche·30m ago
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
0
AM1010101·1h ago
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.

On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.

If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.

0
ChickeNES·1h ago
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
0
s3p·1h ago
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
0
skeledrew·7m ago
Can't wait for it to get to the point where it's like an internet subscription: unlimited tokens 24/7/365 at a low, fixed monthly price.
0
s314·1h ago
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
0
zozbot234·1h ago
> In the Artificial Analysis Intelligence Index

lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.

0
swingboy·45m ago
Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.
0
alasano·1h ago
I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
0
pookieinc·1h ago
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
0
radial_symmetry·1h ago
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
0
WinstonSmith84·1h ago
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
0
jrflo·1h ago
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
0
lanthissa·1h ago
cutting edge fable is for them not you and they're not going to share the metrics until they give you access.
0
bpodgursky·1h ago
Fable 5.5 probably drops soon so it would just be confusing.
0
solenoid0937·1h ago
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.

At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.

0
rfgplk·1h ago
Astra is still the uncontested #1 code generator.
0
solenoid0937·1h ago
Astra is amazing, I love it.
0
dude250711·58m ago
Yeah, especially coupled with Opus for alternative reviews. A massive token burn though.
0
ricardobeat·1h ago
I mean, they worked really hard for this. Back in February everybody loved them.
0
solenoid0937·1h ago
I think all the positive people have just stopped commenting.

The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.

0
taurath·1h ago
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
0
nicoburns·1h ago
5 was definitely bad. 5.5 seems a lot better so far. But still not close to Fable in terms of quality.
0
croemer·1h ago
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
0
ghoshbishakh·1h ago
So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.

In xhigh effort it is a lot cheaper and possibly lot less impressive?

0
pavitheran·1h ago
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
0
mydreamof·56m ago
Cost are bigger than Opus 5.5 for that effort
0
takerofnaps·1h ago
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
0
bayesianbot·1h ago
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
0
square_usual·1h ago
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
0
simianwords·1h ago
Important to note that lower model + higher reasoning gives a different (not higher) quality of response than higher model + lower reasoning.

Some tasks are reasoning shaped by nature and you can't just throw a big model at it.

0
rtuin·1h ago
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
0
limsungkee·1h ago
Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
0
_fw·1h ago
I still can’t find a place for Sonnet models, I never have.

I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.

Give me the frontier, or give me the cheapest form of good enough.

0
calumcl·1h ago
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.

Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.

0
EMM_386·1h ago
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.

Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.

0