[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-claude-sonnet-5-pricing-benchmarks-openrouter-en":3,"article-related-claude-sonnet-5-pricing-benchmarks-openrouter-en":30,"series-model-release-08360412-a7dd-47d8-8318-0e2f12a39a4e":77},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":23,"views":27,"created_at":28,"published_at":29,"topic_cluster_id":11},"08360412-a7dd-47d8-8318-0e2f12a39a4e","claude-sonnet-5-pricing-benchmarks-openrouter-en","Claude Sonnet 5 pricing and benchmarks on OpenRouter","\u003Cp data-speakable=\"summary\">OpenRouter lists \u003Ca href=\"\u002Ftag\u002Fclaude\">Claude\u003C\u002Fa> Sonnet 5 at $2 input and $10 output per million tokens.\u003C\u002Fp>\u003Cp>\u003Ca href=\"\u002Ftag\u002Fanthropic\">Anthropic\u003C\u002Fa>’s \u003Ca href=\"https:\u002F\u002Fopenrouter.ai\u002Fanthropic\u002Fclaude-sonnet-5\" target=\"_blank\" rel=\"noopener\">Claude Sonnet 5\u003C\u002Fa> arrives with a 1M-token context window, adaptive thinking levels, and pricing that starts at $2 per million input tokens. OpenRouter also shows a $10 per million output-token rate, plus provider data that makes it easier to compare speed, uptime, and real-world routing behavior.\u003C\u002Fp>\u003Ctable>\u003Cthead>\u003Ctr>\u003Cth>Metric\u003C\u002Fth>\u003Cth>Value\u003C\u002Fth>\u003Cth>Why it matters\u003C\u002Fth>\u003C\u002Ftr>\u003C\u002Fthead>\u003Ctbody>\u003Ctr>\u003Ctd>Input price\u003C\u002Ftd>\u003Ctd>$2 \u002F 1M tokens\u003C\u002Ftd>\u003Ctd>Sets the base cost for prompts and retrieved context\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Output price\u003C\u002Ftd>\u003Ctd>$10 \u002F 1M tokens\u003C\u002Ftd>\u003Ctd>Matters most for long answers and agent loops\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Context window\u003C\u002Ftd>\u003Ctd>1M tokens\u003C\u002Ftd>\u003Ctd>Lets teams keep much larger working sets in memory\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Release date\u003C\u002Ftd>\u003Ctd>Jun 30, 2026\u003C\u002Ftd>\u003Ctd>Places the model in Anthropic’s newest Sonnet tier\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Best listed latency\u003C\u002Ftd>\u003Ctd>1.22 s\u003C\u002Ftd>\u003Ctd>Shown by Azure on OpenRouter’s provider table\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Best listed throughput\u003C\u002Ftd>\u003Ctd>70 tok\u002Fs\u003C\u002Ftd>\u003Ctd>Useful for chatty apps and agent workloads\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch2>What OpenRouter says Sonnet 5 is for\u003C\u002Fh2>\u003Cp>OpenRouter describes Sonnet 5 as Anthropic’s most capable Sonnet-class model, aimed at coding, agents, and professional work. That phrasing matters because the model page is not just marketing copy; it is paired with provider data, \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> scores, and live traffic signals that show how the model behaves outside of a lab.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787038382800-3z17.png\" alt=\"Claude Sonnet 5 pricing and benchmarks on OpenRouter\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>The model supports adaptive thinking with selectable reasoning effort levels: low, medium, high, max, and x-high. It also accepts text, image, and file inputs, which makes it a better fit for mixed workflows than a plain text-only model. If you want a quick contrast with other releases, our coverage of \u003Ca href=\"\u002Fnews\u002Fclaude-code-update\" target=\"_blank\" rel=\"noopener\">Claude Code updates\u003C\u002Fa> and \u003Ca href=\"\u002Fnews\u002Fopenrouter-model-routing\" target=\"_blank\" rel=\"noopener\">OpenRouter routing\u003C\u002Fa> shows how these tools are being used in production setups.\u003C\u002Fp>\u003Cul>\u003Cli>1M-token context window for long prompts and large codebases\u003C\u002Fli>\u003Cli>Adaptive thinking levels from low through x-high\u003C\u002Fli>\u003Cli>Text, image, and file inputs\u003C\u002Fli>\u003Cli>Updated tokenizer\u003C\u002Fli>\u003Cli>Real-time cyber safeguards for higher-risk dual-use requests\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That last point is easy to miss, but it matters. OpenRouter says Sonnet 5 includes real-time cyber safeguards that block certain high-risk dual-use activities. For teams building agentic products, that changes how you think about prompt design, tool access, and failure modes.\u003C\u002Fp>\u003Ch2>Pricing is simple, but provider costs are not\u003C\u002Fh2>\u003Cp>The headline price is straightforward: $2 per million input tokens and $10 per million output tokens. OpenRouter’s provider table shows how the same model can cost differently depending on where it runs, with some providers charging more and some adding better latency or uptime.\u003C\u002Fp>\u003Cp>Here is the useful part for operators: the average price customers actually pay can be lower than the posted rate because of caching and discounts. That means a model page like this is doing two jobs at once. It tells you the sticker price, and it also gives you a sense of the real bill once routing and provider behavior kick in.\u003C\u002Fp>\u003Cblockquote>“OpenRouter is introducing a new pricing and routing model called the OpenRouter API.” — OpenRouter blog, 2024\u003C\u002Fblockquote>\u003Cp>That line from OpenRouter’s own blog captures the point of the platform. You are not buying a single endpoint; you are buying access to multiple providers that can be swapped under the hood when one is slow or down. OpenRouter’s model page says it can route requests using Balanced, Nitro, or Exacto modes, depending on whether you care most about price, speed, or tool-calling accuracy.\u003C\u002Fp>\u003Cp>For teams comparing vendors, the provider list is more useful than a generic spec sheet. It shows that \u003Ca href=\"https:\u002F\u002Fwww.anthropic.com\" target=\"_blank\" rel=\"noopener\">Anthropic\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fcloud.google.com\u002Fvertex-ai\" target=\"_blank\" rel=\"noopener\">Google Vertex AI\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fazure.microsoft.com\u002Fen-us\u002Fproducts\u002Fai-services\u002Fopenai-service\" target=\"_blank\" rel=\"noopener\">Azure\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Faws.amazon.com\u002Fbedrock\u002F\" target=\"_blank\" rel=\"noopener\">Amazon Bedrock\u003C\u002Fa> can all expose the same model with different economics and performance tradeoffs.\u003C\u002Fp>\u003Ch2>The benchmark numbers tell a practical story\u003C\u002Fh2>\u003Cp>OpenRouter’s benchmark section gives Sonnet 5 strong scores on standardized evaluations, including GPQA Diamond and TAU-Bench. The exact values vary by provider, but the spread is narrow enough to suggest that routing choice matters less for raw capability than for cost, latency, and availability.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787038380657-97gq.png\" alt=\"Claude Sonnet 5 pricing and benchmarks on OpenRouter\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>On the latency and throughput table, Azure on OpenRouter shows the best listed latency at 1.22 seconds and 70 tokens per second throughput. Anthropic’s own endpoint sits close behind with 1.68 seconds latency and 56 tokens per second. Those are the kinds of numbers that matter when a product team is deciding whether a model feels instant or sluggish in a live app.\u003C\u002Fp>\u003Cul>\u003Cli>Azure: 1.22 s latency, 70 tok\u002Fs throughput, 99.96% uptime\u003C\u002Fli>\u003Cli>Anthropic: 1.68 s latency, 56 tok\u002Fs throughput, 99.99% uptime\u003C\u002Fli>\u003Cli>Claude Platform on AWS: 2.76 s latency, 48 tok\u002Fs throughput, 99.96% uptime\u003C\u002Fli>\u003Cli>Google Vertex (Global): 2.98 s latency, 60 tok\u002Fs throughput, 99.89% uptime\u003C\u002Fli>\u003Cli>OpenRouter availability over 24 hours: 99.66%\u003C\u002Fli>\u003Cli>Without routing: 92.70%\u003C\u002Fli>\u003C\u002Ful>\u003Cp>The routing gap is the most interesting number in that set. OpenRouter says availability over the last 24 hours was 99.66% overall, while requests without routing came in at 92.70%. That is a big difference if your app depends on consistent responses and you do not want a provider hiccup to become your outage.\u003C\u002Fp>\u003Cp>OpenRouter also reports 100.00% uptime over three days and 99.80% availability over the same period. For teams running agents or support workflows, that is the kind of operational detail that matters more than a glossy model card.\u003C\u002Fp>\u003Ch2>What the traffic data suggests about real use\u003C\u002Fh2>\u003Cp>The public app rankings on OpenRouter are a handy proxy for actual demand. Claude Sonnet 5 appears in workflows like \u003Ca href=\"https:\u002F\u002Fwww.anthropic.com\u002Fclaude-code\" target=\"_blank\" rel=\"noopener\">Claude Code\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fnousresearch.com\u002Fhermes\" target=\"_blank\" rel=\"noopener\">Hermes Agent\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fwww.descript.com\" target=\"_blank\" rel=\"noopener\">Descript\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fwww.openclaw.ai\" target=\"_blank\" rel=\"noopener\">OpenClaw\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Fwww.framer.com\" target=\"_blank\" rel=\"noopener\">Framer\u003C\u002Fa>. Those are very different products, which tells you the model is being used for code, media editing, messaging automation, and website generation.\u003C\u002Fp>\u003Cp>OpenRouter lists 624B tokens for \u003Ca href=\"\u002Fnews\u002Fclaude-code-desktop-ships-inside-one-app-en\">Claude Code\u003C\u002Fa> alone, followed by 528B for Hermes \u003Ca href=\"\u002Ftag\u002Fagent\">Agent\u003C\u002Fa> and 466B for Descript. That is a strong hint that Sonnet 5 is being pulled into agent-heavy and production-heavy workloads rather than casual chat only. If you are building a product with \u003Ca href=\"\u002Ftag\u002Flong-context\">long context\u003C\u002Fa>, tool use, or repeated edits, those usage patterns are the most relevant signal on the page.\u003C\u002Fp>\u003Cp>There is also a practical takeaway for engineering managers. A model can look good on a benchmark table and still be a poor fit if its provider mix is unstable, expensive, or slow in your region. OpenRouter’s page solves that by putting benchmarks, uptime, and provider pricing side by side.\u003C\u002Fp>\u003Ch2>What to do with Sonnet 5 right now\u003C\u002Fh2>\u003Cp>Claude Sonnet 5 looks best as a serious default for teams that need large context, agent workflows, and predictable ops data in one place. If you are evaluating it, do not stop at the headline price; compare the provider table, test the routing modes, and check whether your workload is output-heavy enough for the $10 per million token rate to matter.\u003C\u002Fp>\u003Cp>The next question is simple: does your app need maximum model quality, or does it need the best mix of quality, latency, and routing resilience? For Sonnet 5, OpenRouter gives you enough numbers to answer that with a real benchmark instead of a guess.\u003C\u002Fp>","OpenRouter lists Claude Sonnet 5 at $2 input and $10 output per million tokens, with a 1M context window and strong coding benchmarks.","openrouter.ai","https:\u002F\u002Fopenrouter.ai\u002Fanthropic\u002Fclaude-sonnet-5",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787038382800-3z17.png","model-release","en","f0cba6ba-77cb-425b-a60d-40e8c928ed2b",[17,18,19,20,21,22],"Claude Sonnet 5","OpenRouter","Anthropic","API pricing","benchmarks","routing",[24,25,26],"Sonnet 5 costs $2 per million input tokens and $10 per million output tokens on OpenRouter.","The model supports a 1M-token context window, adaptive thinking, and text, image, and file inputs.","OpenRouter’s routing and provider data show big differences in latency, uptime, and availability.",0,"2026-08-18T07:32:31.416899+00:00","2026-08-18T07:32:31.409+00:00",{"tags":31,"relatedLang":36,"relatedPosts":40},[32,34],{"name":18,"slug":33},"openrouter",{"name":19,"slug":35},"anthropic",{"id":15,"slug":37,"title":38,"language":39},"claude-sonnet-5-pricing-benchmarks-openrouter-zh","Claude Sonnet 5 在 OpenRouter 的定價與基準","zh",[41,47,53,59,65,71],{"id":42,"slug":43,"title":44,"cover_image":45,"image_url":45,"created_at":46,"category":13},"2f7f6f21-84a5-47d0-8ffd-ff024b95c53c","meta-30b-local-model-zuckerberg-manifesto-en","Meta’s 30B local model and Zuckerberg’s AI manifesto","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787011395428-id25.png","2026-08-18T00:02:51.185597+00:00",{"id":48,"slug":49,"title":50,"cover_image":51,"image_url":51,"created_at":52,"category":13},"37acb0da-5597-4dee-b383-b8d9b11dbac7","cognition-40b-valuation-funding-talks-en","Cognition may be eyeing a $40B valuation","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786993372645-jcb3.png","2026-08-17T19:02:28.472095+00:00",{"id":54,"slug":55,"title":56,"cover_image":57,"image_url":57,"created_at":58,"category":13},"1ebdee84-7b23-4d3a-afe1-c9ff20202d24","shieldstral-turns-moderation-policy-into-one-model-en","Shieldstral turns moderation policy into one model","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786975413452-0upt.png","2026-08-17T14:02:54.826984+00:00",{"id":60,"slug":61,"title":62,"cover_image":63,"image_url":63,"created_at":64,"category":13},"2a151c6e-e731-468d-8924-c4ee731edb1e","claude-opus-5-benchmarks-developers-en","Claude Opus 5 Benchmarks for Developers","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786930369734-fq7o.png","2026-08-17T01:32:26.592024+00:00",{"id":66,"slug":67,"title":68,"cover_image":69,"image_url":69,"created_at":70,"category":13},"cd53c895-decb-45ea-879f-9124307e11b6","anthropic-adds-watermarking-across-claude-products-en","Anthropic adds watermarking across Claude products","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786906982602-8537.png","2026-08-16T19:02:39.927+00:00",{"id":72,"slug":73,"title":74,"cover_image":75,"image_url":75,"created_at":76,"category":13},"262a94e0-b7c4-4276-9e30-909e529306c1","anthropic-ipo-talks-skip-valuation-en","Anthropic’s IPO talks skip valuation for now","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786708967564-i643.png","2026-08-14T12:02:28.507445+00:00",[78,83,88,93,98,103,108,113,118,123],{"id":79,"slug":80,"title":81,"created_at":82},"d4cffde7-9b50-4cc7-bb68-8bc9e3b15477","nvidia-rubin-ai-supercomputer-en","NVIDIA Unveils Rubin: A Leap in AI Supercomputing","2026-03-25T16:24:35.155565+00:00",{"id":84,"slug":85,"title":86,"created_at":87},"eab919b9-fbac-4048-89fc-afad6749ccef","google-gemini-ai-innovations-2026-en","Google's AI Leap with Gemini Innovations in 2026","2026-03-25T16:27:18.841838+00:00",{"id":89,"slug":90,"title":91,"created_at":92},"5f5cfc67-3384-4816-a8f6-19e44d90113d","gap-google-gemini-ai-checkout-en","Gap Teams Up with Google Gemini for AI-Driven Checkout","2026-03-25T16:27:46.483272+00:00",{"id":94,"slug":95,"title":96,"created_at":97},"f6d04567-47f6-49ec-804c-52e61ab91225","ai-model-release-wave-march-2026-en","Navigating the AI Model Release Wave of March 2026","2026-03-25T16:28:45.409716+00:00",{"id":99,"slug":100,"title":101,"created_at":102},"895c150c-569e-4fdf-939d-dade785c990e","small-language-models-transform-ai-en","Small Language Models: Llama 3.2 and Phi-3 Transform AI","2026-03-25T16:30:26.688313+00:00",{"id":104,"slug":105,"title":106,"created_at":107},"38eb1d26-d961-4fd3-ae12-9c4089680f5f","midjourney-v8-alpha-features-pricing-en","Midjourney V8 Alpha: A Deep Dive into Its Features and Pricing","2026-03-26T01:25:36.387587+00:00",{"id":109,"slug":110,"title":111,"created_at":112},"bf36bb9e-3444-4fb8-ab19-0df6bc9d8271","rag-2026-indispensable-ai-bridge-en","RAG in 2026: The Indispensable AI Bridge","2026-03-26T01:28:34.472046+00:00",{"id":114,"slug":115,"title":116,"created_at":117},"60881d6d-2310-44ef-b1fb-7f98e9dd2f0e","xiaomi-mimo-trio-agents-robots-voice-en","Xiaomi’s MiMo trio targets agents, robots, and voice","2026-03-28T03:05:08.899895+00:00",{"id":119,"slug":120,"title":121,"created_at":122},"f063d8d1-41d1-4de4-8ebc-6c40511b9369","xiaomi-mimo-v2-pro-1t-moe-agents-en","Xiaomi MiMo-V2-Pro: 1T MoE Model for Agents","2026-03-28T03:06:19.238032+00:00",{"id":124,"slug":125,"title":126,"created_at":127},"a1379e9a-6785-4ff5-9b0a-8cff55f8264f","cursor-composer-2-started-from-kimi-en","Cursor’s Composer 2 started from Kimi","2026-03-28T03:11:59.132398+00:00"]