[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-baton-long-horizon-robot-manipulation-en":3,"article-related-baton-long-horizon-robot-manipulation-en":29,"series-research-7d7d6420-88d8-4fc5-bf1c-c37dab7e101d":72},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"7d7d6420-88d8-4fc5-bf1c-c37dab7e101d","baton-long-horizon-robot-manipulation-en","BATON tackles long-horizon robot manipulation","\u003Cp data-speakable=\"summary\">BATON improves long-horizon robot manipulation by exploring subtasks separately and tracking transition state in memory.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 11.6% task success gain on RoboMemArena\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Subtask-level exploration with transition-aware memory\u003C\u002Fli>\u003C\u002Ful>\u003Cp>How do you keep a robot from falling apart when a task spans many steps, many contacts, and many chances for one mistake to poison the rest?\u003C\u002Fp>\u003Cp>This paper argues that the usual long-horizon setup is stacked against you: even if a vision-language-action model can handle individual \u003Ca href=\"\u002Ftag\u002Fskills\">skills\u003C\u002Fa>, chaining them together makes errors compound in ways the policy cannot easily recover from. The result is a system that can look competent in short segments but still fail once those segments have to connect into one continuous job.\u003C\u002Fp>\u003Cp>BATON is the authors’ answer to that failure mode. Instead of treating the whole task as one giant exploration problem, it breaks the problem into subtasks and makes each subtask the unit of exploration. It also adds a memory mechanism that is explicitly aware of transitions, so the system can reason not just about how to finish a step, but about whether the next step can actually start from the state it leaves behind.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>Long-horizon robot manipulation is hard because it is not just a sequence of isolated actions. It is a chain of contact-rich skills, and each skill changes the environment in ways that affect the next one. A grasp, a placement, or a wrist repositioning move can all leave the scene in a state that is technically “successful” for the current subtask but unusable for the next.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787032974941-k0fz.png\" alt=\"BATON tackles long-horizon robot manipulation\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>The abstract says current VLA models increasingly master individual skills, but the chain still breaks. That matters because the failure is not always localized: one bad transition can silently constrain later steps, and once the task gets long enough, the system may no longer know which stage actually caused the breakdown.\u003C\u002Fp>\u003Cp>The paper also calls out a second issue with a common recipe: freeze the VLA, let an \u003Ca href=\"\u002Ftag\u002Fllm\">LLM\u003C\u002Fa> \u003Ca href=\"\u002Ftag\u002Fagent\">agent\u003C\u002Fa> plan in language, use analytic primitives for free-space motion, call the VLA only for contact-rich segments, and write adaptation into language memory. That sounds modular, but the authors argue it breaks in two specific ways when tasks stretch out.\u003C\u002Fp>\u003Cul>\u003Cli>Whole-task exploration is expensive, because the cost grows multiplicatively with stages.\u003C\u002Fli>\u003Cli>The system lacks a representation of transitions, so a subtask can finish in a form the next subtask cannot use.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>How BATON works in plain English\u003C\u002Fh2>\u003Cp>BATON changes the unit of exploration. Instead of trying to discover a full long-horizon solution by repeatedly rolling out the entire task, it explores each subtask in the cheaper short-horizon regime. Once a subtask is solved, that solution is stored in memory and later composed into a longer trajectory.\u003C\u002Fp>\u003Cp>The paper frames this as a cost shift from multiplicative to additive. In the abstract’s terms, if one stage needs T episodes, a K-stage task does not require something like T^K exploration effort; BATON aims to make it closer to T*K by solving stages separately. Just as important, failures become easier to attribute because they can be tied to a single stage rather than a whole tangled sequence.\u003C\u002Fp>\u003Cp>The second part of BATON is transition-aware memory. Within a subtask, a verifier agent controls when the VLA should be invoked. The VLA is only called after the wrist view confirms that the scene is ready. That is a practical guardrail: don’t spend the expensive contact-rich policy until the preconditions actually look right.\u003C\u002Fp>\u003Cp>Across subtasks, BATON uses two transition mechanisms. A handoff transition restores an entry state that may have been disturbed by residue from the predecessor. A lookahead transition chooses the strategy whose outcome the successor can inherit. In other words, the system is not just storing “what worked,” but “what worked in a form the next step can accept.”\u003C\u002Fp>\u003Cp>One detail that stands out for engineers: no parameters are updated. BATON is presented as a test-time, agentic orchestration and memory approach rather than a training-time model rewrite. That makes it easier to understand as a control layer around an existing VLA stack, not a new base model.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The abstract gives one concrete \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa>: RoboMemArena, described as a long-horizon benchmark. On that benchmark, BATON improves task success by 11.6% and cumulative success by 14.9% over the state of the art.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787032975940-8bru.png\" alt=\"BATON tackles long-horizon robot manipulation\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Those are meaningful gains, but the abstract does not provide the full experimental table, task breakdown, or per-skill results. It also does not give the number of tasks, the exact robot setup, or the baseline names. So the safe reading is that BATON appears to help at the level the paper cares about most: end-to-end completion over long sequences, not just isolated skill execution.\u003C\u002Fp>\u003Cp>The benchmark numbers matter because they match the paper’s diagnosis. If the main failure is bad composition across stages, then a method that improves transition handling and subtask reuse should help most when the horizon gets long. The reported gains suggest that this is not just a theoretical framing change; it translates into better task completion on the benchmark the authors chose.\u003C\u002Fp>\u003Cp>Still, the abstract leaves open several practical questions. How robust is the verifier when the wrist view is ambiguous? How much of the gain comes from better subtask decomposition versus the transition-aware memory itself? And how well would the approach generalize to different robot bodies, sensors, or manipulation domains? The abstract does not answer those points.\u003C\u002Fp>\u003Ch2>Why developers and robotics teams should care\u003C\u002Fh2>\u003Cp>If you build robot stacks, BATON is interesting because it treats long-horizon failure as a systems problem, not just a model-capacity problem. That is a useful mental model for anyone trying to chain together already-strong policies into something that can survive real multi-stage work.\u003C\u002Fp>\u003Cp>The subtask-first approach is also operationally attractive. Short-horizon exploration is cheaper, failures are easier to localize, and memory can be structured around reusable solutions instead of one-off trajectories. That suggests a path for teams who want to improve reliability without retraining the entire policy stack.\u003C\u002Fp>\u003Cp>The transition-aware part is probably the most practical takeaway. In real manipulation, “success” is often not enough; the next step needs the right pose, the right clearance, or the right object state. BATON’s handoff and lookahead ideas are a reminder that interfaces between skills deserve as much design attention as the skills themselves.\u003C\u002Fp>\u003Cp>At the same time, the paper is not claiming to solve long-horizon manipulation in general. The abstract is explicit that this is a no-parameter-update approach evaluated on one benchmark, and it does not show that BATON removes the underlying difficulty of contact-rich robotics. It shows a cleaner way to manage it.\u003C\u002Fp>\u003Cp>For engineers, that makes BATON worth reading as a pattern: separate exploration by subtask, verify transition readiness, and store state in a way that preserves downstream usability. Even if you never use this exact system, the design principle is broadly applicable to any agentic robotics pipeline that has to survive more than one step at a time.\u003C\u002Fp>\u003Ch2>Bottom line\u003C\u002Fh2>\u003Cp>BATON’s core idea is simple but practical: long-horizon robot manipulation fails when you treat it like one giant problem, so make subtasks the unit of exploration and make transitions explicit in memory.\u003C\u002Fp>\u003Cp>The paper reports 11.6% higher task success and 14.9% higher cumulative success on RoboMemArena, but the abstract does not provide deeper benchmark detail. What it does provide is a clear systems-level argument for how to make multi-stage robot behavior less brittle.\u003C\u002Fp>","BATON improves long-horizon robot manipulation by exploring subtasks separately and tracking transition state in memory.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.16889",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787032974941-k0fz.png","research","en","4a1f6f7d-daeb-47ac-ad84-545e2ab390eb",[17,18,19,20,21],"robot manipulation","vision-language-action","long-horizon planning","memory","transition handling",[23,24,25],"BATON shifts exploration from whole-task rollouts to subtask-level search.","It adds transition-aware memory so one step leaves the right state for the next.","On RoboMemArena, it reports 11.6% higher task success and 14.9% higher cumulative success.",0,"2026-08-18T06:02:26.635592+00:00","2026-08-18T06:02:26.626+00:00",{"tags":30,"relatedLang":31,"relatedPosts":35},[],{"id":15,"slug":32,"title":33,"language":34},"baton-long-horizon-robot-manipulation-zh","BATON 讓長程機器人更穩","zh",[36,42,48,54,60,66],{"id":37,"slug":38,"title":39,"cover_image":40,"image_url":40,"created_at":41,"category":13},"7a79ef84-f0ae-498b-8540-286d89541841","turboquant-new-baseline-long-context-inference-en","TurboQuant is not a niche trick; it is the new baseline for long-cont…","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787041969874-5u8x.png","2026-08-18T08:32:21.653598+00:00",{"id":43,"slug":44,"title":45,"cover_image":46,"image_url":46,"created_at":47,"category":13},"a2592c81-fa86-4316-b223-c31f66e4424d","matrix-multiplication-bound-alphaevolve-en","New matrix-multiplication bound via AlphaEvolve","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787036573596-j6rj.png","2026-08-18T07:02:26.074287+00:00",{"id":49,"slug":50,"title":51,"cover_image":52,"image_url":52,"created_at":53,"category":13},"7b4d1968-cdfa-4998-aa3b-b4c004572661","qvirl-bayesian-irl-uncertainty-en","QVIRL learns rewards with uncertainty","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1787034777458-q78g.png","2026-08-18T06:32:28.062097+00:00",{"id":55,"slug":56,"title":57,"cover_image":58,"image_url":58,"created_at":59,"category":13},"52b5bc33-08cb-4ddc-a272-898e56c6dedf","handover-in-context-learning-state-en","How to hand off LLM session state","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786950182521-034z.png","2026-08-17T07:02:36.188221+00:00",{"id":61,"slug":62,"title":63,"cover_image":64,"image_url":64,"created_at":65,"category":13},"f28b65ab-1e05-4f30-b6f4-e5f4b381f072","marionette-world-state-geometry-appearance-en","Marionette splits game world state from appearance","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786948381354-qx93.png","2026-08-17T06:32:33.803571+00:00",{"id":67,"slug":68,"title":69,"cover_image":70,"image_url":70,"created_at":71,"category":13},"8dd81b24-1d6e-488c-b54b-7aa7dacaa53f","uncertainty-aware-ai-prehistoric-hand-stencils-en","Uncertainty-Aware AI Reads Prehistoric Hand Stencils","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786946587729-mo56.png","2026-08-17T06:02:39.504759+00:00",[73,78,83,88,93,98,103,108,113,118],{"id":74,"slug":75,"title":76,"created_at":77},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":79,"slug":80,"title":81,"created_at":82},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":84,"slug":85,"title":86,"created_at":87},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":89,"slug":90,"title":91,"created_at":92},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":94,"slug":95,"title":96,"created_at":97},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":99,"slug":100,"title":101,"created_at":102},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":104,"slug":105,"title":106,"created_at":107},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":109,"slug":110,"title":111,"created_at":112},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":114,"slug":115,"title":116,"created_at":117},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":119,"slug":120,"title":121,"created_at":122},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]