[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-kimi-k3-test-time-scaling-rules-en":3,"article-related-kimi-k3-test-time-scaling-rules-en":30,"series-industry-e746ff12-ac66-4bc2-9492-a33e2f921677":76},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":23,"views":27,"created_at":28,"published_at":29,"topic_cluster_id":11},"e746ff12-ac66-4bc2-9492-a33e2f921677","kimi-k3-test-time-scaling-rules-en","Kimi K3 maps the new rules of test-time scaling","\u003Cp>What does Kimi K3 say about where frontier LLM scaling is headed?\u003C\u002Fp>\u003Cp data-speakable=\"summary\">Kimi K3 frames test-time compute as the new center of frontier LLM scaling.\u003C\u002Fp>\u003Ctable>\u003Cthead>\u003Ctr>\u003Cth>Item\u003C\u002Fth>\u003Cth>Core idea\u003C\u002Fth>\u003Cth>Scaling mode\u003C\u002Fth>\u003C\u002Ftr>\u003C\u002Fthead>\u003Ctbody>\u003Ctr>\u003Ctd>OpenAI o series\u003C\u002Ftd>\u003Ctd>Reinforcement learning plus test-time reasoning\u003C\u002Ftd>\u003Ctd>Longer inference\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Anthropic extended thinking models\u003C\u002Ftd>\u003Ctd>Adaptive thinking budgets with tool use\u003C\u002Ftd>\u003Ctd>Dynamic inference\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>DeepSeek-R1\u003C\u002Ftd>\u003Ctd>Large-scale RL from a strong pretrained base\u003C\u002Ftd>\u003Ctd>Reasoning emergence\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Kimi K1.5\u003C\u002Ftd>\u003Ctd>RL-driven complex reasoning behavior\u003C\u002Ftd>\u003Ctd>Reasoning emergence\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Kimi K2.5 Agent Swarm\u003C\u002Ftd>\u003Ctd>Parallel agent coordination at test time\u003C\u002Ftd>\u003Ctd>Multi-agent inference\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch2>1. OpenAI o series\u003C\u002Fh2>\u003Cp>The \u003Ca href=\"https:\u002F\u002Fopenai.com\u002F\">OpenAI\u003C\u002Fa> o series is the clearest sign that scaling no longer means only bigger pretraining runs. In this view, test-time compute becomes a second axis of progress, with \u003Ca href=\"\u002Ftag\u002Freinforcement-learning\">reinforcement learning\u003C\u002Fa> used to improve reasoning after training is done.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785718974890-fnrs.png\" alt=\"Kimi K3 maps the new rules of test-time scaling\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>That matters because it changes how capability is bought. Instead of spending only on a larger checkpoint, the model can spend more thinking budget at \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa> time when the task needs it. The result is a system that can trade latency for better answers on harder prompts.\u003C\u002Fp>\u003Cul>\u003Cli>Focus: reinforcement learning for reasoning\u003C\u002Fli>\u003Cli>Scaling unit: inference-time compute\u003C\u002Fli>\u003Cli>Practical effect: more deliberate answers on difficult tasks\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>2. Anthropic extended thinking models\u003C\u002Fh2>\u003Cp>\u003Ca href=\"https:\u002F\u002Fwww.anthropic.com\u002F\">Anthropic\u003C\u002Fa> pushed the idea further by making the thinking budget adaptive. Rather than using the same amount of compute for every request, these models can allocate more or less effort based on the problem in front of them.\u003C\u002Fp>\u003Cp>The other key move is that reasoning and tool use are tied together. That makes the model less like a pure text generator and more like a planner that can decide when to think, when to call a tool, and when to answer directly.\u003C\u002Fp>\u003Cul>\u003Cli>Adaptive thinking budget instead of fixed effort\u003C\u002Fli>\u003Cli>Tool use integrated with reasoning\u003C\u002Fli>\u003Cli>Better fit for tasks with uneven difficulty\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>3. DeepSeek-R1\u003C\u002Fh2>\u003Cp>\u003Ca href=\"https:\u002F\u002Fwww.deepseek.com\u002F\">DeepSeek-R1\u003C\u002Fa> showed that large-scale RL can pull complex reasoning out of a strong pretrained base. The point is not just that the model got better, but that the training recipe made reasoning behaviors emerge more clearly under pressure from reward signals.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785718972831-sdjj.png\" alt=\"Kimi K3 maps the new rules of test-time scaling\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>This is important for the broader scaling story because it suggests that raw pretraining is only part of the equation. If the base model is strong enough, then post-training can reshape it into a much more capable reasoner without changing the core architecture.\u003C\u002Fp>\u003Cul>\u003Cli>Large-scale reinforcement learning\u003C\u002Fli>\u003Cli>Strong pretrained foundation\u003C\u002Fli>\u003Cli>Reasoning gains through post-training\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>4. Kimi K1.5\u003C\u002Fh2>\u003Cp>\u003Ca href=\"https:\u002F\u002Fwww.kimi.com\u002F\">Kimi\u003C\u002Fa> K1.5 sits in the same family of ideas, but it reinforces the claim from another angle. It shows that complex reasoning behavior can be activated through large-scale RL, not only by making the model bigger before deployment.\u003C\u002Fp>\u003Cp>That makes K1.5 a useful marker in the timeline. It links the older scaling logic, which centered on pretraining size, with the newer logic, which treats post-training and inference-time effort as first-class levers for capability.\u003C\u002Fp>\u003Cul>\u003Cli>Complex reasoning from RL\u003C\u002Fli>\u003Cli>Post-training as a capability driver\u003C\u002Fli>\u003Cli>Bridge between pretraining and test-time scaling\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>5. Kimi K2.5 Agent Swarm\u003C\u002Fh2>\u003Cp>Kimi K2.5 \u003Ca href=\"\u002Ftag\u002Fagent\">Agent\u003C\u002Fa> Swarm extends the idea beyond single-model reasoning into parallel coordination. Instead of one chain of thought doing all the work, multiple agents can cooperate at test time, which turns scaling into a coordination problem as much as a compute problem.\u003C\u002Fp>\u003Cp>This is the most forward-looking part of the story. It suggests that the next step after longer reasoning is distributed reasoning, where the system gets better not just by thinking harder, but by thinking together.\u003C\u002Fp>\u003Cul>\u003Cli>Parallel agent collaboration\u003C\u002Fli>\u003Cli>Test-time scaling beyond serial reasoning\u003C\u002Fli>\u003Cli>Useful for tasks that benefit from division of labor\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>How to decide\u003C\u002Fh2>\u003Cp>If you care about the broad direction of frontier LLM research, the main lesson is simple: scaling is no longer only about bigger pretraining runs. The most important gains now come from test-time compute, adaptive reasoning budgets, RL-driven post-training, and multi-agent coordination.\u003C\u002Fp>\u003Cp>Pick the \u003Ca href=\"\u002Ftag\u002Fopenai\">OpenAI\u003C\u002Fa> and \u003Ca href=\"\u002Ftag\u002Fanthropic\">Anthropic\u003C\u002Fa> examples if you want the clearest picture of inference-time reasoning. Pick DeepSeek-R1 and Kimi K1.5 if you want to understand how RL reshapes a pretrained base. Pick Kimi K2.5 Agent Swarm if you want the strongest hint about where the next wave of scaling may go.\u003C\u002Fp>","4 model families show how test-time compute, RL, and agent swarms are reshaping frontier LLM scaling.","zhuanlan.zhihu.com","https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F2065855401434800906",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785718974890-fnrs.png","industry","en","6e506134-abee-4415-b7d7-0b455f9a5dbc",[17,18,19,20,21,22],"Kimi K3","test-time scaling","LLM reasoning","reinforcement learning","agent swarms","frontier models",[24,25,26],"Test-time compute is now a core scaling axis, not an add-on.","RL can unlock complex reasoning from strong pretrained models.","Parallel agent coordination points to the next stage of inference scaling.",1,"2026-08-03T01:02:30.065567+00:00","2026-08-03T01:02:30.054+00:00",{"tags":31,"relatedLang":36,"relatedPosts":40},[32,34],{"name":20,"slug":33},"reinforcement-learning",{"name":19,"slug":35},"llm-reasoning",{"id":15,"slug":37,"title":38,"language":39},"kimi-k3-report-test-time-scaling-4-directions-zh","Kimi K3 點出測試時擴展的4條路","zh",[41,46,52,58,64,70],{"id":42,"slug":43,"title":44,"cover_image":11,"image_url":11,"created_at":45,"category":13},"2d727daf-559d-44b3-bc2b-3ff53137afef","ai-weekly-2026-w32-en","AI Weekly: 2026-07-27 ~ 2026-08-03","2026-08-03T04:00:39.96111+00:00",{"id":47,"slug":48,"title":49,"cover_image":50,"image_url":50,"created_at":51,"category":13},"61ebc42d-aa1c-43c0-9d6f-3ce43f714971","salp-liquidation-ai-trades-revealed-en","What the SALP liquidation reveals about AI trades","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785717176689-5kgk.png","2026-08-03T00:32:33.409316+00:00",{"id":53,"slug":54,"title":55,"cover_image":56,"image_url":56,"created_at":57,"category":13},"b3199b83-9b2b-45f3-ba99-8ee782d45130","claude-security-test-became-real-breach-en","Claude’s security test became a real breach","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785715402243-s100.png","2026-08-03T00:02:55.308879+00:00",{"id":59,"slug":60,"title":61,"cover_image":62,"image_url":62,"created_at":63,"category":13},"7c78bbe7-bdbd-461c-b536-90b37dd24ac1","x-posts-let-execs-shape-the-ai-story-en","X posts let execs shape the AI story","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785697388377-hlfd.png","2026-08-02T19:02:41.347626+00:00",{"id":65,"slug":66,"title":67,"cover_image":68,"image_url":68,"created_at":69,"category":13},"3302d464-a3d5-4550-b328-4b2c7c1a89b7","jensen-huang-agi-definition-lowers-the-bar-en","Jensen Huang’s AGI definition lowers the bar","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785695567157-abd0.png","2026-08-02T18:32:24.259106+00:00",{"id":71,"slug":72,"title":73,"cover_image":74,"image_url":74,"created_at":75,"category":13},"017693e6-8d4a-409e-a848-5bdf697ee08d","claude-2026-limit-changes-capacity-story-en","Claude’s 2026 limit changes are a capacity story, not a product story","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785693771298-z5r3.png","2026-08-02T18:02:25.374847+00:00",[77,82,87,92,97,102,107,112,117,122],{"id":78,"slug":79,"title":80,"created_at":81},"d35a1bd9-e709-412e-a2df-392df1dc572a","ai-impact-2026-developments-market-en","AI's Impact in 2026: Key Developments and Market Shifts","2026-03-25T16:20:33.205823+00:00",{"id":83,"slug":84,"title":85,"created_at":86},"5ed27921-5fd6-492e-8c59-78393bf37710","trumps-ai-legislative-framework-en","Trump's AI Legislative Framework: What's Inside?","2026-03-25T16:22:20.005325+00:00",{"id":88,"slug":89,"title":90,"created_at":91},"e454a642-f03c-4794-b185-5f651aebbaca","nvidia-gtc-2026-key-highlights-innovations-en","NVIDIA GTC 2026: Key Highlights and Innovations","2026-03-25T16:22:47.882615+00:00",{"id":93,"slug":94,"title":95,"created_at":96},"0ebb5b16-774a-4922-945d-5f2ce1df5a6d","claude-usage-diversifies-learning-curves-en","Claude Usage Diversifies, Learning Curves Emerge","2026-03-25T16:25:50.770376+00:00",{"id":98,"slug":99,"title":100,"created_at":101},"69934e86-2fc5-4280-8223-7b917a48ace8","openclaw-ai-commoditization-concerns-en","OpenClaw's Rise Raises Concerns of AI Model Commoditization","2026-03-25T16:26:30.582047+00:00",{"id":103,"slug":104,"title":105,"created_at":106},"b4b2575b-2ac8-46b2-b90e-ab1d7c060797","google-gemini-ai-rollout-2026-en","Google's Gemini AI Rollout Extended to 2026","2026-03-25T16:28:14.808842+00:00",{"id":108,"slug":109,"title":110,"created_at":111},"6e18bc65-42ae-4ad0-b564-67d7f66b979e","meta-llama4-fabricated-results-scandal-en","Meta's Llama 4 Scandal: Fabricated AI Test Results Unveiled","2026-03-25T16:29:15.482836+00:00",{"id":113,"slug":114,"title":115,"created_at":116},"bf888e9d-08be-4f47-996c-7b24b5ab3500","accenture-mistral-ai-deployment-en","Accenture and Mistral AI Team Up for AI Deployment","2026-03-25T16:31:01.894655+00:00",{"id":118,"slug":119,"title":120,"created_at":121},"5382b536-fad2-49c6-ac85-9eb2bae49f35","mistral-ai-high-stakes-2026-en","Mistral AI: Facing High Stakes in 2026","2026-03-25T16:31:39.941974+00:00",{"id":123,"slug":124,"title":125,"created_at":126},"9da3d2d6-b669-4971-ba1d-17fdb3548ed5","cursors-meteoric-rise-pressures-en","Cursor's Meteoric Rise Faces Industry Pressures","2026-03-25T16:32:21.899217+00:00"]