[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-coderescue-budget-calibrated-recovery-routing-en":3,"article-related-coderescue-budget-calibrated-recovery-routing-en":30,"series-research-302ac5a7-8d8f-462e-88ea-739f7aa89fb1":73},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":29},"302ac5a7-8d8f-462e-88ea-739f7aa89fb1","coderescue-budget-calibrated-recovery-routing-en","CodeRescue routes coding-agent recovery by budget","\u003Cp data-speakable=\"summary\">CodeRescue learns when \u003Ca href=\"\u002Fnews\u002Fkimi-k3-benchmark-evaluation-guide-coding-agents-en\">coding agents\u003C\u002Fa> should keep recovering cheaply or escalate under a budget.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 35% of mean recovery cost\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Supervised recovery router plus CRC budget calibration\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Coding agents do not just fail; they often fail with useful execution feedback. That changes the control problem from a simple “pick a model” decision into a post-failure routing problem: should the \u003Ca href=\"\u002Ftag\u002Fagent\">agent\u003C\u002Fa> spend a little more cheap compute, or should it hand the case to a stronger model?\u003C\u002Fp>\u003Cp>This paper argues that the usual cost-aware playbook is too blunt for that setting. Instead of treating failure as a one-way escalation trigger, \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.19338\">CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents\u003C\u002Fa> models recovery as a choice among heterogeneous actions, then adds a calibration layer so the same router can be deployed under different budgets without retraining.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>The core issue is familiar to anyone building coding agents: the first attempt is rarely the whole story. If the code runs, you get signals from the runtime, tests, or errors. Those signals can make another cheap attempt worthwhile, even when the initial answer was wrong.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784703782443-v9xu.png\" alt=\"CodeRescue routes coding-agent recovery by budget\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Most existing cost-aware systems, the paper says, use a cascade pattern: start with a cheap model, and if it fails, escalate the hard case to a stronger and more expensive model. That works when failure is just a dead end. It works less well when failure is informative and cheap recovery can still succeed.\u003C\u002Fp>\u003Cp>So the paper reframes the decision after a failed attempt. The question is not only “is this case hard?” but “given the failure feedback, which recovery action gives the best tradeoff between cost and success?” That is the routing problem CodeRescue tries to solve.\u003C\u002Fp>\u003Ch2>How the method works in plain English\u003C\u002Fh2>\u003Cp>The method has two layers. First, the authors formulate post-failure decisions as recovery routing over heterogeneous actions. In practice, that means the router is choosing between different recovery paths rather than making a single binary cheap-versus-expensive call.\u003C\u002Fp>\u003Cp>They train that router in a supervised way from execution rollouts. The abstract does not spell out every feature or model detail, but the key idea is straightforward: use the history of failed runs and their outcomes to learn which recovery action tends to work in which situation.\u003C\u002Fp>\u003Cp>Then comes the calibration layer. The paper adds Conformal Risk Control, or CRC, to select a deployment-time cost penalty without retraining. That matters because budgets change. A router that only works for one fixed cost setting is awkward to operate; a router with a tunable penalty can be reused across deployment regimes.\u003C\u002Fp>\u003Cp>The CRC layer is also described as providing marginal expected-cost control under exchangeability. In plain terms, the authors are trying to make the budget knob statistically meaningful, not just heuristic. That gives operators a way to shift the system toward cheaper or more aggressive behavior without rebuilding the router from scratch.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The evaluation uses held-out failures from five coding benchmarks. The abstract does not list the \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> names or their individual scores, so there is no full per-benchmark breakdown here. What it does say is that cheap recovery and escalation show complementary success patterns across those failures.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784703777848-p0vd.png\" alt=\"CodeRescue routes coding-agent recovery by budget\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>That complementarity is the important empirical point. It suggests the choice is not simply “cheap is worse” or “expensive is always safer.” Some failures are worth another low-cost recovery pass, while others should be escalated. CodeRescue is designed to separate those cases instead of forcing a one-size-fits-all cascade.\u003C\u002Fp>\u003Cp>The paper says the calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline. The strongest concrete number in the abstract is this: in the main GPT-5.4-nano\u002FGPT-5.4 setting, one CRC-calibrated frontier point exceeds the always-escalate solve rate while using 35% of its mean recovery cost.\u003C\u002Fp>\u003Cp>That is a useful result for practitioners because it means the system is not just trading away accuracy for cost. At least at one operating point, it can beat the always-escalate strategy on solve rate while spending much less on recovery. The abstract does not provide the full frontier curve, so we cannot say how stable that advantage is across all budgets.\u003C\u002Fp>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you are building a coding agent, the practical lesson is that failure handling is a budgeted control problem, not just a fallback mechanism. A failed execution can be valuable signal, and the best next move may be another cheap recovery step rather than an immediate jump to a larger model.\u003C\u002Fp>\u003Cp>That matters for both product and infrastructure teams. On the product side, it can improve solve rate without forcing every hard case onto the most expensive path. On the infrastructure side, it gives you a more explicit way to manage \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa> spend after failures, which is where many agent workflows quietly burn budget.\u003C\u002Fp>\u003Cp>The calibration angle is also useful. Deployment budgets are rarely static, and retraining a router every time the cost target changes is operationally annoying. A CRC-based penalty selector is a cleaner control surface if the assumptions hold in your environment.\u003C\u002Fp>\u003Ch2>Limits and open questions\u003C\u002Fh2>\u003Cp>The abstract is clear about the high-level method, but it leaves out a lot of implementation detail. We do not get the exact router architecture, the supervised labels used for training, the specific coding benchmarks, or the shape of the recovered actions. Those details matter if you want to reproduce the system or compare it to your own agent stack.\u003C\u002Fp>\u003Cp>There is also an important statistical caveat. CRC is described as providing marginal expected-cost control under exchangeability. That is a useful guarantee, but it depends on assumptions that may not hold cleanly when your production workload shifts, your toolchain changes, or your test distribution drifts.\u003C\u002Fp>\u003Cp>Finally, the abstract reports one strong operating point, not a complete deployment story. We know one calibrated frontier point beats always-escalate solve rate at 35% of its mean recovery cost, but we do not know the latency impact, the failure modes, or how sensitive the system is to different coding tasks and model pairs.\u003C\u002Fp>\u003Ch2>The bottom line\u003C\u002Fh2>\u003Cp>CodeRescue is trying to make coding agents smarter after the first failure. Instead of escalating everything, it learns when cheap recovery is still worth trying and then adds calibration so the same policy can adapt to different budgets.\u003C\u002Fp>\u003Cp>For engineers, the takeaway is simple: post-failure routing is a real optimization problem, and this paper shows that a calibrated router can improve the cost-quality tradeoff over fixed actions and a binary cascade. The main result is promising, but the abstract leaves enough unanswered that the real test will be how well it generalizes outside the reported benchmarks.\u003C\u002Fp>\u003Cul>\u003Cli>It treats failed code execution as useful feedback, not just a dead end.\u003C\u002Fli>\u003Cli>It combines supervised recovery routing with Conformal Risk Control for budget tuning.\u003C\u002Fli>\u003Cli>It reports one frontier point that beats always-escalate solve rate at 35% of mean recovery cost.\u003C\u002Fli>\u003C\u002Ful>","CodeRescue learns when coding agents should keep recovering cheaply or escalate under a budget.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.19338",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784703782443-v9xu.png","research","en","a2a16ea8-c295-4a4e-a887-89294bd40f74",[17,18,19,20,21],"coding agents","budget control","conformal risk control","recovery routing","cost-aware inference",[23,24,25],"Failed executions can guide another cheap recovery attempt instead of immediate escalation.","CodeRescue pairs a supervised router with CRC so budgets can change without retraining.","The reported frontier includes a point that beats always-escalate solve rate at 35% cost.",1,"2026-07-22T07:02:33.432859+00:00","2026-07-22T07:02:33.423+00:00","42d8dc4d-43f7-4279-be44-62ff24e64864",{"tags":31,"relatedLang":32,"relatedPosts":36},[],{"id":15,"slug":33,"title":34,"language":35},"coderescue-budget-calibrated-recovery-routing-zh","CodeRescue 用預算路由修復代理","zh",[37,43,49,55,61,67],{"id":38,"slug":39,"title":40,"cover_image":41,"image_url":41,"created_at":42,"category":13},"370eab09-3a2b-44cb-8900-2ef2fa2687de","appearance-pointers-region-control-dits-en","Appearance Pointers bring region control to DiTs","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784701977102-6s2p.png","2026-07-22T06:32:28.561668+00:00",{"id":44,"slug":45,"title":46,"cover_image":47,"image_url":47,"created_at":48,"category":13},"5df4c442-0663-4423-b917-00de6965f627","gear-cuts-copying-long-context-reasoning-en","GEAR cuts copying in long-context reasoning","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784700183820-l53a.png","2026-07-22T06:02:30.227905+00:00",{"id":50,"slug":51,"title":52,"cover_image":53,"image_url":53,"created_at":54,"category":13},"84f909d5-e578-49ad-9f8e-c8cafc7562ea","rag17-sod1-als-nature-medicine-template-en","RAG-17 turns SOD1-ALS data into a template","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784678589143-aopn.png","2026-07-22T00:02:48.345502+00:00",{"id":56,"slug":57,"title":58,"cover_image":59,"image_url":59,"created_at":60,"category":13},"33248bb8-c831-4d24-a0e5-b8cc13cac750","survey-of-large-language-models-en","A Survey of Large Language Models","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784629987559-3qtb.png","2026-07-21T10:32:29.824097+00:00",{"id":62,"slug":63,"title":64,"cover_image":65,"image_url":65,"created_at":66,"category":13},"332f5dcb-3420-4277-9ac9-4cb3e690c3c7","evaluating-memory-in-llm-agents-en","How to test memory in LLM agents","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784628193788-ty9w.png","2026-07-21T10:02:36.648611+00:00",{"id":68,"slug":69,"title":70,"cover_image":71,"image_url":71,"created_at":72,"category":13},"4cccdf92-dbaf-4ec3-9ef2-cc2a4e8a1a13","persona-steering-llm-capabilities-analysis-en","How persona steering changes LLM behavior","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1784626384361-j5on.png","2026-07-21T09:32:28.472784+00:00",[74,79,84,89,94,99,104,109,114,119],{"id":75,"slug":76,"title":77,"created_at":78},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":80,"slug":81,"title":82,"created_at":83},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":85,"slug":86,"title":87,"created_at":88},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":90,"slug":91,"title":92,"created_at":93},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":95,"slug":96,"title":97,"created_at":98},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":100,"slug":101,"title":102,"created_at":103},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":105,"slug":106,"title":107,"created_at":108},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":110,"slug":111,"title":112,"created_at":113},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":115,"slug":116,"title":117,"created_at":118},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":120,"slug":121,"title":122,"created_at":123},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]