Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · September 28, 2026 · concept

Claude Sonnet 5.5Claude Sonnet 5.5 System Cardfrontier-modelsagent-securityclaude-codeagent-harness

Claude Sonnet 5.5 Got Better at Cyber and Worse at Saying No

What the Sonnet 5.5 system card means for anyone running the new default model inside their own agent harness

Two numbers in the Claude Sonnet 5.5 system card point in opposite directions, and almost nobody reading the launch post will see both. On a 10-challenge subset of CyScenarioBench, a test of whether a model can plan and run multi-stage cyber operations, Sonnet 5.5 completed 46.1% of challenges. Sonnet 5 completed 0.7%.

On a 112-task evaluation of malicious computer use, the same model refused 79.46% of the time. Sonnet 5 refused 84.68%.

So the model got dramatically more capable at the dangerous thing and a little less likely to decline the dangerous thing, and it did both at exactly the same price as the model it replaces. That combination is the story of this release for anyone who builds agents, and it has a practical answer that has nothing to do with waiting for the next model.

Why this release lands differently

Price slots decide where models end up. Anthropic released Claude Sonnet 5.5 on September 28 at $2 per million input tokens and $10 per million output, identical to Sonnet 5, and says it "typically needs far fewer tokens to do the same work," costing "up to 30% less per task." It scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, and lands two points below Opus 5.5 on GDPval-AA.

Those are vendor numbers. Treat them that way until your own suite agrees.

Even discounted, the direction is clear. The mid-tier model, the one teams wire into CI jobs, support bots, overnight refactors and browser agents because it is affordable, now carries capability that used to cost more. Mid-tier models run in the most places with the least supervision. Whatever safety properties the model carries get tested at that scale, in harnesses Anthropic never sees.

My position is simple. If you run Sonnet 5.5 outside Anthropic's own apps, treat its refusal behavior as a bonus you cannot budget on, and build the no into your environment.

Read the two tables together

The launch post leads with capability. The system card is where the tradeoff shows, and it helps to know what each number measures.

The malicious computer-use evaluation gives the model GUI and command-line tools in a sandbox and asks it to do harmful work across three areas: surveillance and unauthorized data collection, generation and distribution of harmful content, and scaled abuse. The table heading is explicit that these are "results without mitigations." That makes 79.46% a property of the model alone, before any classifier touches the request. The card offers no explanation for the drop. It notes only that Sonnet 5.5 "performs below Sonnet 5 and Mythos 5.1 but has an identical refusal rate to our most recent model, Opus 5.5."

Five points sounds small. Read it as a miss rate instead: Sonnet 5 failed to refuse about 15% of those tasks, Sonnet 5.5 about 21%. On a fleet of agents running unattended, that is the number that compounds.

The cyber numbers explain why the miss rate matters more now than it did last week. Beyond the CyScenarioBench jump, the card reports ExploitBench results of 11.53 flags captured on average in its AutoNudge setup and complete arbitrary code execution exploits in 43.4% of runs, and calls Sonnet 5.5 "a significant step up in cyber capabilities from Claude Sonnet 5" while still short of Opus 5.5. A model that refuses slightly less often is a minor concern when it cannot finish the job. It is a different concern when it can.

Where the no actually lives now

Anthropic's answer to the gap is not in the weights. It is in a classifier system wrapped around the model. For cyber misuse, the card describes the same three stages Opus 5.5 uses: "a probe that looks at Claude's internal activations, a lightweight classifier that runs on Claude Sonnet 5.5 itself, and a trained LLM classifier." The key sentence about what happens when those stages fire is this one:

"On most interfaces, Claude Sonnet 5.5 falls back to Claude Sonnet 5 for requests that are blocked by our cyber classifier system. This happens automatically in our own applications; on the API, the developer must opt in to automatic fallbacks."

Read the second half twice. In the Claude apps, a blocked cyber request gets handed to the older, less capable model without you doing anything. On the API, where most agent harnesses live, that fallback is something you have to turn on.

This is a reasonable design. Security researchers need these capabilities, and the card points them to a "forthcoming updated Cyber Verification Program." Blocking everything would break legitimate work. But it moves the safety property from something the model is to something the deployment does, and the deployment is you as soon as you call the API from your own code.

I think this is the shift worth naming. A refusal rate used to be a rough stand-in for how much a model would help with bad work. For Sonnet 5.5, the honest description is layered: the weights say no about four times in five on this test, Anthropic's classifiers catch more on their own surfaces, and your harness gets whatever you configure.

What got better, and why it points the same way

The card is not a list of regressions. On the question most agent builders care about day to day, the model improved.

Prompt injection, where hostile content in a web page or file hijacks the agent, got harder. On Gray Swan's indirect injection benchmark, attack success was 0.4% at k=1, 2.7% at k=10 and 3.4% at k=15, down from Sonnet 5's 0.7%, 5.1% and 6.7%. In browser use, tested without safeguards, no attack succeeded across 110 scenarios. The card still says Sonnet 5.5 trails Opus 5.5 and Fable 5.1 on the Gray Swan measure.

Put the two results side by side and the picture sharpens. Sonnet 5.5 is harder for a third party to hijack and a little easier for its operator to point at bad work on purpose. The first problem is getting solved in the model. The second one is getting handed to deployment. If you run the deployment, it is handed to you.

Put this into practice

None of this needs new tooling. It needs a few settings and an afternoon.

Turn on the fallback if you use the API. If your harness calls Sonnet 5.5 directly, find the automatic fallback option the card describes and opt in. It costs you nothing on legitimate traffic and restores the behavior Anthropic's own apps get by default.

Set effort explicitly. Claude Code defaults Sonnet 5.5 to Medium effort and the Platform defaults to High, per the launch post. The same prompt can behave differently in each. Pin the effort level in code so your evals and your production runs match. If you ran Sonnet with thinking disabled, the launch post says you need the new between_tools setting before upgrading.

Rerun your own misuse and task suite before swapping. Take the ten worst things your agent could be talked into doing with its real tools (delete a branch, email a customer list, scrape a site you do not own) and run them against both models. You are not trying to reproduce Anthropic's 112 tasks. You are finding your own miss rate.

Put the no in the environment. This is the part that holds up across every model release:

  • Egress allowlists, so a browser or shell agent can reach only the hosts its job needs.
  • Scoped, short-lived credentials instead of a long-lived token with write access to everything.
  • Sandboxes with no path to production data unless a human opens one.
  • An approval step on irreversible actions: sends, deletes, payments, deploys.

A refusal from the model is still welcome. It just stops being the only thing standing between a bad instruction and a bad outcome.

Honest limitations

Several things here cut against my own argument, and you should weigh them.

The refusal drop is five points on one 112-task evaluation, measured by the vendor, with no confidence interval in the summary I read. It could narrow or widen on other task sets. I would not claim Sonnet 5.5 is "unsafe," and neither does the card.

The capability numbers are also vendor-run. A 60-point jump on Terminal-Bench 4.0 is large enough that I would want independent reproduction before quoting it as settled. The CyScenarioBench figure comes from a 10-challenge subset, which is a small sample.

The classifier layer may be doing more than the card spells out. "Most interfaces" is not a full list. The card does not say how Bedrock, Vertex or Azure handle the fallback, and I did not find that documented. If you deploy through a cloud marketplace, ask your provider rather than assume.

And the environment controls I recommend are not free. Egress allowlists break agents that need to browse widely. Approval steps slow down exactly the unattended work that made Sonnet attractive in the first place. That tradeoff is real, and the right balance depends on what your agent can reach.

The call is yours

The easy read of today's launch is that the mid-tier model got much better for the same money. That is true, and for most workloads it is good news.

The fuller read is that the model now carries capability that deserves a firmer boundary than its own judgment. Anthropic built that boundary for its own apps. On the API, it waits for you to switch it on.

So check which side of that line your agents run on. Then decide, deliberately, where your no lives.

Sources: Anthropic, Introducing Claude Sonnet 5.5 · Claude Sonnet 5.5 System Card (PDF)


Medium metadata

  • Title: Claude Sonnet 5.5 Got Better at Cyber and Worse at Saying No
  • Subtitle: What the Sonnet 5.5 system card means for anyone running the new default model inside their own agent harness
  • Tags: Claude, AI Agents, AI Safety, Cybersecurity, Anthropic
  • Canonical: import from the fervorai.dev URL
  • Reading time: about 8 minutes