[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-claude-opus-5-lowest-safety-audit-score-en":3,"article-related-claude-opus-5-lowest-safety-audit-score-en":29,"series-industry-b76b6998-3bc6-4c39-abc4-eb2e8fd750a9":74},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"b76b6998-3bc6-4c39-abc4-eb2e8fd750a9","claude-opus-5-lowest-safety-audit-score-en","Claude Opus 5 posts the lowest safety audit score","\u003Cp>Which \u003Ca href=\"\u002Ftag\u002Fanthropic\">Anthropic\u003C\u002Fa> model scored best in the new behavior audit?\u003C\u002Fp>\u003Cp data-speakable=\"summary\">\u003Ca href=\"\u002Fnews\u002Fclaude-opus-5-undercuts-fable-5-price-en\">Claude Opus\u003C\u002Fa> 5 posted the lowest score in Anthropic’s automated behavior audit.\u003C\u002Fp>\u003Ctable>\u003Cthead>\u003Ctr>\u003Cth>Model\u003C\u002Fth>\u003Cth>Audit score\u003C\u002Fth>\u003Cth>Rank among these 4\u003C\u002Fth>\u003C\u002Ftr>\u003C\u002Fthead>\u003Ctbody>\u003Ctr>\u003Ctd>Claude Opus 5\u003C\u002Ftd>\u003Ctd>2.30\u003C\u002Ftd>\u003Ctd>1\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Claude Opus 4.8\u003C\u002Ftd>\u003Ctd>2.85\u003C\u002Ftd>\u003Ctd>2\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Claude Mythos 5\u003C\u002Ftd>\u003Ctd>2.81\u003C\u002Ftd>\u003Ctd>3\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Claude Sonnet 5\u003C\u002Ftd>\u003Ctd>3.35\u003C\u002Ftd>\u003Ctd>4\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch2>1. Claude Opus 5\u003C\u002Fh2>\u003Cp>\u003Ca href=\"https:\u002F\u002Fwww.anthropic.com\u002F\">Anthropic\u003C\u002Fa> says this model got a score of 2.30 in its automated behavior audit, which is the lowest among the four recent models mentioned here. In this audit, lower is better because the test is designed to catch deception, risky behavior, and susceptibility to malicious prompting.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785283367536-dqkm.png\" alt=\"Claude Opus 5 posts the lowest safety audit score\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cul>\u003Cli>Audit score: 2.30\u003C\u002Fli>\u003Cli>Best result in the group\u003C\u002Fli>\u003Cli>Focused on behavior under pressure\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That makes \u003Ca href=\"\u002Ftag\u002Fclaude\">Claude\u003C\u002Fa> Opus 5 the model to watch if you care more about safer conduct than raw capability claims. The number does not tell the full story about quality, but it does give a direct comparison point against the other recent Anthropic releases.\u003C\u002Fp>\u003Ch2>2. Claude Opus 4.8\u003C\u002Fh2>\u003Cp>Claude Opus 4.8 came in at 2.85, which is higher than Opus 5 but still better than Claude Sonnet 5 in this audit. For readers comparing model generations, this gives a simple signal that the newer Opus release improved on the same behavior metric.\u003C\u002Fp>\u003Cul>\u003Cli>Audit score: 2.85\u003C\u002Fli>\u003Cli>Below Sonnet 5\u003C\u002Fli>\u003Cli>Older than Opus 5\u003C\u002Fli>\u003C\u002Ful>\u003Cp>If you are tracking whether Anthropic’s updates are moving in the right direction on safety-related behavior, Opus 4.8 is a useful midpoint. It is not the top performer in this set, but it helps show the gap between the latest release and the prior version.\u003C\u002Fp>\u003Ch2>3. Claude Mythos 5\u003C\u002Fh2>\u003Cp>\u003Ca href=\"\u002Ftag\u002Fclaude-mythos\">Claude Mythos\u003C\u002Fa> 5 scored 2.81, placing it just above Claude Opus 4.8 and just below Claude Opus 5. That narrow spread suggests these recent models are fairly close on this particular audit, even if the order still matters.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785283366446-v0jf.png\" alt=\"Claude Opus 5 posts the lowest safety audit score\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cul>\u003Cli>Audit score: 2.81\u003C\u002Fli>\u003Cli>Near Opus 4.8\u003C\u002Fli>\u003Cli>Closer to the top than Sonnet 5\u003C\u002Fli>\u003C\u002Ful>\u003Cp>For a reader trying to understand the relative positioning, Mythos 5 is the model that sits in the middle of the pack. It is not the lowest score, but it is clearly ahead of the weakest performer in this comparison.\u003C\u002Fp>\u003Ch2>4. Claude Sonnet 5\u003C\u002Fh2>\u003Cp>Claude Sonnet 5 recorded a 3.35 score, which is the highest number in the set and therefore the weakest result on this audit. Because the audit is built to penalize deception, harmful behavior, and being manipulated into bad actions, a higher score is less desirable here.\u003C\u002Fp>\u003Cul>\u003Cli>Audit score: 3.35\u003C\u002Fli>\u003Cli>Highest score in this group\u003C\u002Fli>\u003Cli>Least favorable result on this test\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That does not mean Sonnet 5 is broadly bad, only that it did worse on this specific behavior check than the other three models listed. If your main concern is this kind of safety audit, it is the one to compare most carefully against the lower-scoring options.\u003C\u002Fp>\u003Ch2>How to decide\u003C\u002Fh2>\u003Cp>If you want the single best result from this audit, Claude Opus 5 is the clear pick because it posted the lowest score. If you want to compare generations or product tiers, Opus 4.8 and Mythos 5 give you the middle range, while Sonnet 5 shows the upper end of the scores in this group.\u003C\u002Fp>\u003Cp>For safety-focused readers, the key point is simple: this audit rewards lower numbers, and Anthropic’s latest Opus model led the set. For everyone else, the table is the fastest way to see how close the recent models are and where the biggest gap sits.\u003C\u002Fp>","4 recent Anthropic models were audited for deceptive or harmful behavior, and Claude Opus 5 scored the lowest at 2.30.","zhuanlan.zhihu.com","https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F2064290264823379497",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785283367536-dqkm.png","industry","en","29f74ed1-6812-470d-b4fb-88c0383fae6c",[17,18,19,20,21],"Claude Opus 5","Anthropic","behavior audit","model safety","AI evaluation",[23,24,25],"Claude Opus 5 had the lowest audit score at 2.30.","The audit measures deceptive, risky, and maliciously influenced behavior.","Claude Sonnet 5 had the highest score at 3.35 in this group.",0,"2026-07-29T00:02:27.011313+00:00","2026-07-29T00:02:26.999+00:00",{"tags":30,"relatedLang":33,"relatedPosts":37},[31],{"name":18,"slug":32},"anthropic",{"id":15,"slug":34,"title":35,"language":36},"claude-opus-5-behavior-audit-lowest-score-zh","Claude Opus 5 在行为审计里垫底？先看 4 款模型分数","zh",[38,44,50,56,62,68],{"id":39,"slug":40,"title":41,"cover_image":42,"image_url":42,"created_at":43,"category":13},"c04e81bc-6d2f-4674-8d67-98c908874850","risc-v-is-a-real-platform-now-en","RISC-V is past the hobby phase and should be treated as a real platfo…","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785294170094-xdj0.png","2026-07-29T03:02:23.765577+00:00",{"id":45,"slug":46,"title":47,"cover_image":48,"image_url":48,"created_at":49,"category":13},"4915d6ba-d302-4df4-be11-6c7aa25897b7","nvidia-openai-250b-ai-backstop-talks-en","Nvidia and OpenAI discuss a $250B AI backstop","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785286988479-w45c.png","2026-07-29T01:02:30.786732+00:00",{"id":51,"slug":52,"title":53,"cover_image":54,"image_url":54,"created_at":55,"category":13},"ed220d9a-5fe4-4792-82e1-05e3261f5613","anthropic-cognizant-claude-enterprise-expansion-en","Anthropic expands Claude partnership with Cognizant","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785240177666-4wmu.png","2026-07-28T12:02:33.507322+00:00",{"id":57,"slug":58,"title":59,"cover_image":60,"image_url":60,"created_at":61,"category":13},"1f0bc490-1b40-4495-bba8-6f5c323883a0","windsurf-cascade-agentic-ide-workflow-en","Windsurf’s Cascade turns IDE edits into agent work","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785225821346-zga7.png","2026-07-28T08:03:10.208737+00:00",{"id":63,"slug":64,"title":65,"cover_image":66,"image_url":66,"created_at":67,"category":13},"cd208923-8ab6-43ca-975a-ed851bc9b8ab","fuerteventura-2026-freestyle-crowned-two-winners-en","Fuerteventura 2026 freestyle crowned two clear winners","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785223971552-44bl.png","2026-07-28T07:32:20.378369+00:00",{"id":69,"slug":70,"title":71,"cover_image":72,"image_url":72,"created_at":73,"category":13},"cec74680-40e3-44b0-8c48-219843e409f5","musk-grok-imagine-odyssey-film-bet-en","Musk’s Grok Imagine bets big on an Odyssey film","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785200570572-hhcu.png","2026-07-28T01:02:25.549951+00:00",[75,80,85,90,95,100,105,110,115,120],{"id":76,"slug":77,"title":78,"created_at":79},"d35a1bd9-e709-412e-a2df-392df1dc572a","ai-impact-2026-developments-market-en","AI's Impact in 2026: Key Developments and Market Shifts","2026-03-25T16:20:33.205823+00:00",{"id":81,"slug":82,"title":83,"created_at":84},"5ed27921-5fd6-492e-8c59-78393bf37710","trumps-ai-legislative-framework-en","Trump's AI Legislative Framework: What's Inside?","2026-03-25T16:22:20.005325+00:00",{"id":86,"slug":87,"title":88,"created_at":89},"e454a642-f03c-4794-b185-5f651aebbaca","nvidia-gtc-2026-key-highlights-innovations-en","NVIDIA GTC 2026: Key Highlights and Innovations","2026-03-25T16:22:47.882615+00:00",{"id":91,"slug":92,"title":93,"created_at":94},"0ebb5b16-774a-4922-945d-5f2ce1df5a6d","claude-usage-diversifies-learning-curves-en","Claude Usage Diversifies, Learning Curves Emerge","2026-03-25T16:25:50.770376+00:00",{"id":96,"slug":97,"title":98,"created_at":99},"69934e86-2fc5-4280-8223-7b917a48ace8","openclaw-ai-commoditization-concerns-en","OpenClaw's Rise Raises Concerns of AI Model Commoditization","2026-03-25T16:26:30.582047+00:00",{"id":101,"slug":102,"title":103,"created_at":104},"b4b2575b-2ac8-46b2-b90e-ab1d7c060797","google-gemini-ai-rollout-2026-en","Google's Gemini AI Rollout Extended to 2026","2026-03-25T16:28:14.808842+00:00",{"id":106,"slug":107,"title":108,"created_at":109},"6e18bc65-42ae-4ad0-b564-67d7f66b979e","meta-llama4-fabricated-results-scandal-en","Meta's Llama 4 Scandal: Fabricated AI Test Results Unveiled","2026-03-25T16:29:15.482836+00:00",{"id":111,"slug":112,"title":113,"created_at":114},"bf888e9d-08be-4f47-996c-7b24b5ab3500","accenture-mistral-ai-deployment-en","Accenture and Mistral AI Team Up for AI Deployment","2026-03-25T16:31:01.894655+00:00",{"id":116,"slug":117,"title":118,"created_at":119},"5382b536-fad2-49c6-ac85-9eb2bae49f35","mistral-ai-high-stakes-2026-en","Mistral AI: Facing High Stakes in 2026","2026-03-25T16:31:39.941974+00:00",{"id":121,"slug":122,"title":123,"created_at":124},"9da3d2d6-b669-4971-ba1d-17fdb3548ed5","cursors-meteoric-rise-pressures-en","Cursor's Meteoric Rise Faces Industry Pressures","2026-03-25T16:32:21.899217+00:00"]