[MODEL] 7 min readOraCore Editors

Claude Sonnet 5 pricing and benchmarks on OpenRouter

OpenRouter lists Claude Sonnet 5 at $2 input and $10 output per million tokens, with a 1M context window and strong coding benchmarks.

Share LinkedIn
Claude Sonnet 5 pricing and benchmarks on OpenRouter

OpenRouter lists Claude Sonnet 5 at $2 input and $10 output per million tokens.

Anthropic’s Claude Sonnet 5 arrives with a 1M-token context window, adaptive thinking levels, and pricing that starts at $2 per million input tokens. OpenRouter also shows a $10 per million output-token rate, plus provider data that makes it easier to compare speed, uptime, and real-world routing behavior.

MetricValueWhy it matters
Input price$2 / 1M tokensSets the base cost for prompts and retrieved context
Output price$10 / 1M tokensMatters most for long answers and agent loops
Context window1M tokensLets teams keep much larger working sets in memory
Release dateJun 30, 2026Places the model in Anthropic’s newest Sonnet tier
Best listed latency1.22 sShown by Azure on OpenRouter’s provider table
Best listed throughput70 tok/sUseful for chatty apps and agent workloads

What OpenRouter says Sonnet 5 is for

Get the latest AI news in your inbox

Weekly picks of model releases, tools, and deep dives — no spam, unsubscribe anytime.

No spam. Unsubscribe at any time.

OpenRouter describes Sonnet 5 as Anthropic’s most capable Sonnet-class model, aimed at coding, agents, and professional work. That phrasing matters because the model page is not just marketing copy; it is paired with provider data, benchmark scores, and live traffic signals that show how the model behaves outside of a lab.

Claude Sonnet 5 pricing and benchmarks on OpenRouter

The model supports adaptive thinking with selectable reasoning effort levels: low, medium, high, max, and x-high. It also accepts text, image, and file inputs, which makes it a better fit for mixed workflows than a plain text-only model. If you want a quick contrast with other releases, our coverage of Claude Code updates and OpenRouter routing shows how these tools are being used in production setups.

  • 1M-token context window for long prompts and large codebases
  • Adaptive thinking levels from low through x-high
  • Text, image, and file inputs
  • Updated tokenizer
  • Real-time cyber safeguards for higher-risk dual-use requests

That last point is easy to miss, but it matters. OpenRouter says Sonnet 5 includes real-time cyber safeguards that block certain high-risk dual-use activities. For teams building agentic products, that changes how you think about prompt design, tool access, and failure modes.

Pricing is simple, but provider costs are not

The headline price is straightforward: $2 per million input tokens and $10 per million output tokens. OpenRouter’s provider table shows how the same model can cost differently depending on where it runs, with some providers charging more and some adding better latency or uptime.

Here is the useful part for operators: the average price customers actually pay can be lower than the posted rate because of caching and discounts. That means a model page like this is doing two jobs at once. It tells you the sticker price, and it also gives you a sense of the real bill once routing and provider behavior kick in.

“OpenRouter is introducing a new pricing and routing model called the OpenRouter API.” — OpenRouter blog, 2024

That line from OpenRouter’s own blog captures the point of the platform. You are not buying a single endpoint; you are buying access to multiple providers that can be swapped under the hood when one is slow or down. OpenRouter’s model page says it can route requests using Balanced, Nitro, or Exacto modes, depending on whether you care most about price, speed, or tool-calling accuracy.

For teams comparing vendors, the provider list is more useful than a generic spec sheet. It shows that Anthropic, Google Vertex AI, Azure, and Amazon Bedrock can all expose the same model with different economics and performance tradeoffs.

The benchmark numbers tell a practical story

OpenRouter’s benchmark section gives Sonnet 5 strong scores on standardized evaluations, including GPQA Diamond and TAU-Bench. The exact values vary by provider, but the spread is narrow enough to suggest that routing choice matters less for raw capability than for cost, latency, and availability.

Claude Sonnet 5 pricing and benchmarks on OpenRouter

On the latency and throughput table, Azure on OpenRouter shows the best listed latency at 1.22 seconds and 70 tokens per second throughput. Anthropic’s own endpoint sits close behind with 1.68 seconds latency and 56 tokens per second. Those are the kinds of numbers that matter when a product team is deciding whether a model feels instant or sluggish in a live app.

  • Azure: 1.22 s latency, 70 tok/s throughput, 99.96% uptime
  • Anthropic: 1.68 s latency, 56 tok/s throughput, 99.99% uptime
  • Claude Platform on AWS: 2.76 s latency, 48 tok/s throughput, 99.96% uptime
  • Google Vertex (Global): 2.98 s latency, 60 tok/s throughput, 99.89% uptime
  • OpenRouter availability over 24 hours: 99.66%
  • Without routing: 92.70%

The routing gap is the most interesting number in that set. OpenRouter says availability over the last 24 hours was 99.66% overall, while requests without routing came in at 92.70%. That is a big difference if your app depends on consistent responses and you do not want a provider hiccup to become your outage.

OpenRouter also reports 100.00% uptime over three days and 99.80% availability over the same period. For teams running agents or support workflows, that is the kind of operational detail that matters more than a glossy model card.

What the traffic data suggests about real use

The public app rankings on OpenRouter are a handy proxy for actual demand. Claude Sonnet 5 appears in workflows like Claude Code, Hermes Agent, Descript, OpenClaw, and Framer. Those are very different products, which tells you the model is being used for code, media editing, messaging automation, and website generation.

OpenRouter lists 624B tokens for Claude Code alone, followed by 528B for Hermes Agent and 466B for Descript. That is a strong hint that Sonnet 5 is being pulled into agent-heavy and production-heavy workloads rather than casual chat only. If you are building a product with long context, tool use, or repeated edits, those usage patterns are the most relevant signal on the page.

There is also a practical takeaway for engineering managers. A model can look good on a benchmark table and still be a poor fit if its provider mix is unstable, expensive, or slow in your region. OpenRouter’s page solves that by putting benchmarks, uptime, and provider pricing side by side.

What to do with Sonnet 5 right now

Claude Sonnet 5 looks best as a serious default for teams that need large context, agent workflows, and predictable ops data in one place. If you are evaluating it, do not stop at the headline price; compare the provider table, test the routing modes, and check whether your workload is output-heavy enough for the $10 per million token rate to matter.

The next question is simple: does your app need maximum model quality, or does it need the best mix of quality, latency, and routing resilience? For Sonnet 5, OpenRouter gives you enough numbers to answer that with a real benchmark instead of a guess.