[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-surgical-wam-video-pretraining-robot-control-en":3,"article-related-surgical-wam-video-pretraining-robot-control-en":29,"series-research-4c94994e-d58b-4f24-a480-ad026fd60e04":74},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"4c94994e-d58b-4f24-a480-ad026fd60e04","surgical-wam-video-pretraining-robot-control-en","Surgical WAM uses video to train robot control","\u003Cp data-speakable=\"summary\">Surgical WAM shows that action-free video pretraining can improve closed-loop surgical robot control with limited demonstrations.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: Average success rate improved from 63.5% to 77.8%\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Jointly predicts future endoscopic frames and executable action chunks\u003C\u002Fli>\u003C\u002Ful>\u003Cp>How can \u003Ca href=\"\u002Fnews\u002Ffaitheyes-tool-faithful-vision-agents-en\">you train\u003C\u002Fa> a surgical robot when synchronized action labels are scarce and expensive? This paper argues that the answer is to squeeze more value out of abundant endoscopic video, then use that visual prior to make better use of a fixed budget of teleoperated demonstrations.\u003C\u002Fp>\u003Cp>The practical appeal is straightforward: surgical manipulation is hard because it needs precise contact handling, long-horizon planning, and bimanual coordination, but collecting the kind of data that usually powers robot learning is slow and costly. If video-only pretraining can transfer into real closed-loop control, that changes the economics of surgical robot learning.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>The paper starts from a familiar robotics bottleneck: action-labeled demonstrations are the scarce resource. In surgical settings, teleoperated trajectories with synchronized kinematics are expensive to gather, while endoscopic video is much easier to collect.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786514579882-zhmr.png\" alt=\"Surgical WAM uses video to train robot control\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>That imbalance matters because most learning pipelines want both what the robot saw and what it did. Without action labels, video can help with perception or simulation-style modeling, but it often stops short of improving the actual controller. The authors say that gap is exactly where surgical world models have been underused.\u003C\u002Fp>\u003Cp>So the central question is not whether video is useful in general. It is whether action-free video pretraining improves closed-loop surgical manipulation when the action-labeled budget is fixed.\u003C\u002Fp>\u003Ch2>How Surgical WAM works in plain English\u003C\u002Fh2>\u003Cp>The proposed system is called Surgical WAM, short for Surgical World-Action Model. It is described as a unified generative model built on Cosmos Policy.\u003C\u002Fp>\u003Cp>Instead of treating video modeling and control as separate problems, Surgical WAM learns to predict two things at once: future endoscopic observations and executable surgical robot action chunks. That makes the model more than a passive predictor of scene dynamics. It is trained to produce action sequences that can actually drive the robot.\u003C\u002Fp>\u003Cp>The training recipe has two stages. First, the model learns surgical visual dynamics from action-free video. Then it is fine-tuned on the limited set of action-labeled demonstrations available under the fixed budget.\u003C\u002Fp>\u003Cp>At deployment time, Surgical WAM runs as a closed-loop, receding-horizon controller. It executes a short prefix of each predicted action chunk, observes the result, and replans from the new observation. In other words, it does not blindly commit to a long action plan; it keeps updating as the scene changes.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The evaluation is on a suite of four simulated surgical manipulation tasks. The paper reports that video pretraining raises the average success rate from 63.5% to 77.8%.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786514581706-04as.png\" alt=\"Surgical WAM uses video to train robot control\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>That is an absolute gain of 14.3 percentage points, and the abstract highlights a 20-point improvement on PegTransfer. The largest gains show up on contact-rich and bimanual tasks, which is a useful signal because those are exactly the kinds of behaviors where short-horizon visual cues and interaction dynamics matter most.\u003C\u002Fp>\u003Cp>Those are the only concrete \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> numbers given in the abstract, so there is no broader performance table to lean on here. Still, the direction of the result is clear: pretraining on action-free video appears to give the controller a better starting point than learning only from the limited action-labeled set.\u003C\u002Fp>\u003Cul>\u003Cli>Average success rate: 63.5% → 77.8%\u003C\u002Fli>\u003Cli>PegTransfer: +20 percentage points\u003C\u002Fli>\u003Cli>Best gains: contact-rich and bimanual tasks\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Why developers and robotics teams should care\u003C\u002Fh2>\u003Cp>For engineers, the interesting part is not just that the model predicts video. It is that the learned visual dynamics are being used as a control prior. That is a more direct path from cheap observational data to robot behavior than using video only for offline analysis.\u003C\u002Fp>\u003Cp>If this pattern holds beyond the simulated tasks in the paper, it suggests a practical scaling strategy for surgical robotics: collect lots of unlabeled endoscopic video, then spend your limited action-labeled budget where it matters most. That could reduce dependence on large teleoperation datasets, which are the hard part of surgical robot learning.\u003C\u002Fp>\u003Cp>The receding-horizon setup is also a sensible engineering choice for a domain where conditions can shift quickly. By replanning after executing only a short prefix of each predicted chunk, the controller can react to what actually happened instead of trusting a long open-loop rollout.\u003C\u002Fp>\u003Ch2>Limitations and open questions\u003C\u002Fh2>\u003Cp>The abstract only reports results in simulation, so the paper does not yet show real operating-room deployment or hardware validation on a physical surgical robot. That is an important boundary for anyone thinking about translation.\u003C\u002Fp>\u003Cp>It also does not provide benchmark details beyond the four simulated tasks named in the abstract, so it is hard to judge how broad the gains are outside that setup. The paper’s claim is narrower and more defensible: under a fixed action-label budget, action-free video pretraining helps closed-loop control in the tested simulated tasks.\u003C\u002Fp>\u003Cp>Another open question is how far the approach generalizes across surgical procedures, camera setups, and robot platforms. The abstract positions action-free video pretraining as a practical path forward, but the real test will be whether the same recipe keeps working as the environment gets messier and the task distribution changes.\u003C\u002Fp>\u003Ch2>Bottom line\u003C\u002Fh2>\u003Cp>Surgical WAM makes a clear case that abundant endoscopic video is not just passive context; it can become a training signal for control when paired with a world-action model. For teams building surgical robotics systems, the takeaway is simple: if action labels are the bottleneck, video pretraining may be the cheapest way to push performance forward.\u003C\u002Fp>","Surgical WAM shows that action-free video pretraining can improve closed-loop surgical robot control with limited demonstrations.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.11204",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786514579882-zhmr.png","research","en","51473b63-b17b-492a-8dc0-b12069a19b49",[17,18,19,20,21],"surgical robotics","world models","video pretraining","closed-loop control","endoscopic video",[23,24,25],"Action-free endoscopic video can improve surgical robot control when demonstrations are limited.","Surgical WAM predicts both future observations and executable action chunks in one model.","The reported gains come from simulation, so real-world validation is still an open question.",1,"2026-08-12T06:02:32.763172+00:00","2026-08-12T06:02:32.747+00:00",{"tags":30,"relatedLang":33,"relatedPosts":37},[31],{"name":18,"slug":32},"world-models",{"id":15,"slug":34,"title":35,"language":36},"surgical-wam-video-pretraining-robot-control-zh","Surgical WAM 用影片訓練手術機器人控制","zh",[38,44,50,56,62,68],{"id":39,"slug":40,"title":41,"cover_image":42,"image_url":42,"created_at":43,"category":13},"b400fb5d-3c21-4a6a-8383-988225159548","sparse-autoencoders-set-level-instability-en","Sparse Autoencoders Don’t Behave Like Feature Bags","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786518181029-jyb4.png","2026-08-12T07:02:32.315345+00:00",{"id":45,"slug":46,"title":47,"cover_image":48,"image_url":48,"created_at":49,"category":13},"605dd415-e62d-455a-bb4b-d1d2fa487c1b","convawg-controlled-vawg-dialogue-generation-en","ConVAWG generates controlled VAWG dialogues","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786516377348-lc87.png","2026-08-12T06:32:30.834532+00:00",{"id":51,"slug":52,"title":53,"cover_image":54,"image_url":54,"created_at":55,"category":13},"6d197f27-628f-4a63-883d-81a0d9f5c4b5","swe-bench-verified-model-leaderboard-limit-en","SWE-bench Verified has stopped being a clean model leaderboard","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786498365866-w6ew.png","2026-08-12T01:32:20.055621+00:00",{"id":57,"slug":58,"title":59,"cover_image":60,"image_url":60,"created_at":61,"category":13},"30d3b27a-e5fb-4c3c-aed2-b9f078b8be23","dutch-government-llm-benchmark-values-en","Dutch Government LLMs Need More Than Accuracy","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786431775018-wj70.png","2026-08-11T07:02:26.380783+00:00",{"id":63,"slug":64,"title":65,"cover_image":66,"image_url":66,"created_at":67,"category":13},"999ef42b-6d1d-40b3-bdf0-dc87c93f3c89","mmdiff-multimodal-feature-discovery-control-en","MMDiff maps and steers multimodal features","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786429983548-5x3u.png","2026-08-11T06:32:31.358093+00:00",{"id":69,"slug":70,"title":71,"cover_image":72,"image_url":72,"created_at":73,"category":13},"8bb7a700-6170-4000-9902-a24f78586cca","tts-evaluators-miss-more-than-naturalness-en","TTS evaluators miss more than naturalness","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786428179624-mu0v.png","2026-08-11T06:02:29.309089+00:00",[75,80,85,90,95,100,105,110,115,120],{"id":76,"slug":77,"title":78,"created_at":79},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":81,"slug":82,"title":83,"created_at":84},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":86,"slug":87,"title":88,"created_at":89},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":91,"slug":92,"title":93,"created_at":94},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":96,"slug":97,"title":98,"created_at":99},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":101,"slug":102,"title":103,"created_at":104},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":106,"slug":107,"title":108,"created_at":109},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":111,"slug":112,"title":113,"created_at":114},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":116,"slug":117,"title":118,"created_at":119},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":121,"slug":122,"title":123,"created_at":124},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]