[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-humantracker-human-aligned-motion-tracking-benchmark-en":3,"article-related-humantracker-human-aligned-motion-tracking-benchmark-en":29,"series-research-a9281570-db5d-4f07-82e5-5546cd022691":73},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"a9281570-db5d-4f07-82e5-5546cd022691","humantracker-human-aligned-motion-tracking-benchmark-en","HumanTracker fixes humanoid motion eval blind spots","\u003Cp data-speakable=\"summary\">153 hours of motion data and HumanScore make humanoid tracking evaluation better match human judgment.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 153 hours of optical motion trajectories\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Preference-aligned metric trained on human motion-pair judgments\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Humanoid motion tracking sounds straightforward until you try to evaluate it in a way that matches what people actually notice. A tracker can look fine under average pose error and still fail in the ways that matter most in video: unstable support, foot skating, and touch-down timing that feels wrong.\u003C\u002Fp>\u003Cp>\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13555\">HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark\u003C\u002Fa> is trying to close that gap. The paper argues that the standard kinematic metrics used in humanoid tracking are useful, but incomplete, and that the usual test suites are too small and too narrow to stress contact-heavy, long-horizon motion.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>The core issue is evaluation mismatch. In humanoid tracking, a system can score well on per-frame pose differences while still producing motion that looks wrong to a person watching the video. That happens because kinematic error measures geometry, not the physical artifacts that make motion feel broken.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786690977850-fb9v.png\" alt=\"HumanTracker fixes humanoid motion eval blind spots\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>The abstract calls out two recurring failure modes: unstable support and incorrect contacts. In practical terms, that means a robot or avatar may appear to slide a foot, miss the moment a foot should land, or otherwise lose the sense of grounded motion even if the joint angles are close frame by frame.\u003C\u002Fp>\u003Cp>The second problem is dataset scale and diversity. The paper says widely used test suites are small and do not cover enough motion variety to stress contact-rich behavior over long horizons. For developers, that matters because a \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> that only tests easy or repetitive motion can hide the kinds of failures that show up once a system is deployed.\u003C\u002Fp>\u003Ch2>What HumanTracker contains\u003C\u002Fh2>\u003Cp>HumanTracker is presented as a benchmark designed to be both perceptually aligned and scalable. The benchmark contains approximately 153 hours of optical motion trajectories collected from multiple professional performers. That is a meaningful jump in coverage relative to the small test suites the abstract criticizes, though the abstract does not provide a direct size comparison against existing benchmarks.\u003C\u002Fp>\u003Cp>The motions are organized into four motion families, and the benchmark includes text labels for fine-grained diagnosis. That detail matters because it suggests the benchmark is not just a leaderboard dataset; it is also meant to help users understand where a tracker fails, not only whether it fails.\u003C\u002Fp>\u003Cp>For engineers, the text labels are especially useful as an analysis layer. If a system struggles with a certain family of motion, you want to know whether the issue is contact timing, balance, a particular style of movement, or something else. The abstract does not spell out the four families, so the paper itself would be the place to look for the exact taxonomy.\u003C\u002Fp>\u003Ch2>How HumanScore works in plain English\u003C\u002Fh2>\u003Cp>The other major contribution is HumanScore, described as a preference-aligned metric trained on 12K motion pairs containing 24K motions. In plain English, that means the authors did not just define a new formula from geometry alone; they trained a metric to predict which motion humans prefer when comparing two candidates.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786690976265-o71j.png\" alt=\"HumanTracker fixes humanoid motion eval blind spots\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>This is a useful shift because it tries to encode the kinds of judgments people naturally make when watching motion: Does the movement feel stable? Do the contacts happen at the right time? Does the motion look physically credible, not just numerically close?\u003C\u002Fp>\u003Cp>The abstract does not describe the model architecture, training objective, or labeling protocol in detail, so we should not overstate how general this metric is. But the basic idea is clear: align evaluation with human preference rather than relying only on kinematic error.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The abstract reports that, across representative state-of-the-art trackers, HumanScore better predicts human preferences and exposes contact and stability failures that kinematic metrics often miss. That is the paper’s central result, and it is the kind of result that matters more in practice than a small gain in a single numeric score.\u003C\u002Fp>\u003Cp>What the abstract does not include are explicit benchmark numbers for those tracker comparisons. There are no reported percentages, rank changes, or absolute score tables in the provided text, so any stronger claim would be speculation. The safe takeaway is that the authors claim improved alignment with human judgment, not a specific measured margin.\u003C\u002Fp>\u003Cp>That limitation matters. Without the full paper, we do not know how large the preference-prediction gain is, how robust it is across motion families, or how much label noise may be present in the pairwise judgments. We also do not know whether HumanScore generalizes beyond the benchmark settings used in the study.\u003C\u002Fp>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you build humanoid policies, teleoperation systems, or whole-body imitation pipelines, this paper points to a familiar failure mode: optimizing the wrong metric. A system can look good in training logs and still produce motion that humans find obviously broken.\u003C\u002Fp>\u003Cp>HumanTracker is useful because it treats evaluation as a product problem, not just a math problem. The benchmark is designed to surface the kinds of errors that users notice, and HumanScore is designed to rank outputs in a way that better matches human preference. That can help teams debug models, compare trackers more honestly, and avoid overfitting to metrics that miss contact quality.\u003C\u002Fp>\u003Cp>There is also a broader lesson here for anyone building embodied systems: if the output is meant to look natural to people, the evaluation should reflect perception, not just geometry. The paper is making that case for humanoid motion tracking specifically, but the same tension shows up in other generative and control tasks too.\u003C\u002Fp>\u003Ch2>Limits and open questions\u003C\u002Fh2>\u003Cp>The abstract gives a strong high-level pitch, but it leaves several practical questions unanswered. We do not get the benchmark’s exact motion-family definitions, the annotation process behind the text labels, or the details of how HumanScore was trained and validated.\u003C\u002Fp>\u003Cp>We also do not know from the abstract how expensive it is to use HumanScore compared with standard kinematic metrics, whether it requires extra model \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa>, or how sensitive it is to new motion styles outside the benchmark distribution. Those details will matter if teams want to adopt it in a real evaluation pipeline.\u003C\u002Fp>\u003Cul>\u003Cli>HumanTracker expands evaluation with a large motion benchmark built from professional performers.\u003C\u002Fli>\u003Cli>HumanScore is trained on pairwise preference data, not just pose geometry.\u003C\u002Fli>\u003Cli>The main value is catching contact and stability failures that standard kinematic metrics can miss.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Bottom line: HumanTracker is a benchmark paper with a very practical goal — make humanoid motion evaluation line up better with what humans actually see. For teams working on teleoperation or imitation, that is the difference between optimizing for a spreadsheet and optimizing for believable motion.\u003C\u002Fp>","HumanTracker adds 153 hours of motion data and a preference-aligned metric for humanoid tracking.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13555",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786690977850-fb9v.png","research","en","70584f73-54b3-4548-944b-7c596e1e3db5",[17,18,19,20,21],"humanoid tracking","motion benchmark","preference-aligned metric","human evaluation","robotics",[23,24,25],"HumanTracker adds about 153 hours of motion trajectories with labeled motion families.","HumanScore is trained on 12K motion pairs to better match human preference.","The paper argues standard kinematic metrics miss contact and stability failures.",0,"2026-08-14T07:02:31.016013+00:00","2026-08-14T07:02:31.003+00:00",{"tags":30,"relatedLang":32,"relatedPosts":36},[31],{"name":21,"slug":21},{"id":15,"slug":33,"title":34,"language":35},"humantracker-human-aligned-motion-tracking-benchmark-zh","HumanTracker補上人形評測盲點","zh",[37,43,49,55,61,67],{"id":38,"slug":39,"title":40,"cover_image":41,"image_url":41,"created_at":42,"category":13},"6b0b3b1a-f397-44d3-ad83-207b4e87d877","omni-scientist-full-stack-ai-science-en","OmniScientist aims for full-stack AI science","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786689180533-adlk.png","2026-08-14T06:32:31.952472+00:00",{"id":44,"slug":45,"title":46,"cover_image":47,"image_url":47,"created_at":48,"category":13},"62360ad1-f133-491c-9389-966b8d532d46","autodesign-meta-harness-optimization-posters-en","AutoDesign learns better poster-making harnesses","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786687374399-4oto.png","2026-08-14T06:02:27.082839+00:00",{"id":50,"slug":51,"title":52,"cover_image":53,"image_url":53,"created_at":54,"category":13},"5a953549-e09c-43e6-856c-63c394e85997","test-time-harnesses-weak-model-transfer-en","Test-Time Harnesses Transfer Skills Without Retraining","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786604571709-b71l.png","2026-08-13T07:02:25.968754+00:00",{"id":56,"slug":57,"title":58,"cover_image":59,"image_url":59,"created_at":60,"category":13},"84526e03-6caf-4b7e-a910-64f2b712da66","dreamfly-aerial-vln-memory-planning-en","DreamFly improves aerial VLN with memory and planning","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786602779859-brjh.png","2026-08-13T06:32:24.41046+00:00",{"id":62,"slug":63,"title":64,"cover_image":65,"image_url":65,"created_at":66,"category":13},"84e73ffd-ac52-4e86-88cc-abb15eeb9c5e","ava-encoder-agent-native-video-representation-en","AVA-Encoder turns films into editable knowledge graphs","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786600972460-6yms.png","2026-08-13T06:02:22.174937+00:00",{"id":68,"slug":69,"title":70,"cover_image":71,"image_url":71,"created_at":72,"category":13},"b400fb5d-3c21-4a6a-8383-988225159548","sparse-autoencoders-set-level-instability-en","Sparse Autoencoders Don’t Behave Like Feature Bags","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786518181029-jyb4.png","2026-08-12T07:02:32.315345+00:00",[74,79,84,89,94,99,104,109,114,119],{"id":75,"slug":76,"title":77,"created_at":78},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":80,"slug":81,"title":82,"created_at":83},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":85,"slug":86,"title":87,"created_at":88},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":90,"slug":91,"title":92,"created_at":93},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":95,"slug":96,"title":97,"created_at":98},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":100,"slug":101,"title":102,"created_at":103},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":105,"slug":106,"title":107,"created_at":108},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":110,"slug":111,"title":112,"created_at":113},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":115,"slug":116,"title":117,"created_at":118},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":120,"slug":121,"title":122,"created_at":123},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]