[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-pretrain-q-functions-online-rl-finetuning-en":3,"article-related-pretrain-q-functions-online-rl-finetuning-en":29,"series-research-f234ef6f-2934-4e01-bc50-3132313c0d7a":75},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"f234ef6f-2934-4e01-bc50-3132313c0d7a","pretrain-q-functions-online-rl-finetuning-en","Do You Need to Pretrain Q-Functions?","\u003Cp data-speakable=\"summary\">Online RL fine-tuning can work better when Q-functions start from diverse policy rollouts.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 1.26x average improvement\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Initialization via Policy Ensemble bootstraps Q-learning from pooled rollouts\u003C\u002Fli>\u003C\u002Ful>\u003Cp>If you already have a pretrained policy, the obvious next step in value-based RL seems to be pretraining the Q-function too. This paper argues that instinct is often wrong, and shows a simpler path that can work better in online fine-tuning.\u003C\u002Fp>\u003Cp>That matters because Q-functions are a core piece of many RL systems: they guide action selection, shape exploration, and can make or break stability during adaptation. If the initialization is mismatched, you can spend compute pretraining something that does not actually help the policy you end up with.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>The paper looks at a specific question in the pretrain-then-finetune pipeline: if you start with a pretrained policy, should you also pretrain the Q-function on offline data before online RL fine-tuning?\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785391372788-de5m.png\" alt=\"Do You Need to Pretrain Q-Functions?\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Conventional wisdom says yes. The authors say the evidence is weaker than people assume. In their study, naive Q-function pretraining often gives little benefit over random initialization, even when the downstream goal is to fine-tune a pretrained base policy.\u003C\u002Fp>\u003Cp>The reason, as they frame it, is a mismatch between what the offline phase teaches and what online fine-tuning actually needs. The Q-function learned during pretraining is aligned with the pretrained policy’s behavior, not necessarily with the Q-function that the online process converges to later.\u003C\u002Fp>\u003Ch2>Why naive pretraining can miss the mark\u003C\u002Fh2>\u003Cp>In plain English: the Q-function is not just a generic scorekeeper. It is tied to the policy it is evaluating and improving. If the policy changes during online fine-tuning, a Q-function that was optimized around the old behavior may not be the best starting point.\u003C\u002Fp>\u003Cp>The abstract says this gap persists even after offline value maximization. That is an important detail, because it suggests the issue is not simply that the pretraining was weak. The problem is structural: the target you optimize offline is not the same target that matters once fine-tuning starts.\u003C\u002Fp>\u003Cp>For practitioners, that means a “more pretraining” instinct is not automatically safer. In some setups, spending effort on Q pretraining may not buy you much if the learned value landscape is anchored to the wrong policy distribution.\u003C\u002Fp>\u003Ch2>How the method works in plain English\u003C\u002Fh2>\u003Cp>To address the mismatch, the authors propose Initialization via Policy Ensemble, or IPE. The idea is simple: train multiple diverse policies, collect rollouts from all of them, pool that data, and use it to bootstrap Q-function learning for online RL.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785391380478-6eo3.png\" alt=\"Do You Need to Pretrain Q-Functions?\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Instead of initializing the Q-function from a single pretrained policy’s trajectory distribution, IPE starts from a broader set of behaviors. That diversity is the key technical move, because it gives the Q-function a more robust starting point for the online phase.\u003C\u002Fp>\u003Cp>This is not described as a complicated architecture change. It is a training strategy: diversify the policy sources, combine their rollouts, and use that pooled experience to initialize value learning before online fine-tuning begins.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The abstract says the authors systematically study whether Q-function pretraining helps when fine-tuning on top of a pretrained base policy. Their main finding is that naive Q-function pretraining often provides little benefit over random initialization.\u003C\u002Fp>\u003Cp>They then evaluate IPE across a suite of challenging continuous control benchmarks. The reported result is an average 1.26x improvement in fine-tuning performance over naive Q-function pretraining.\u003C\u002Fp>\u003Cp>One thing the abstract does not provide is the full \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> table, task list, or absolute scores. So while the relative gain is clear, readers should not assume this paper proves a universal win across every RL setting.\u003C\u002Fp>\u003Cul>\u003Cli>Naive Q pretraining often underperforms expectations\u003C\u002Fli>\u003Cli>IPE uses multiple diverse policies rather than one source policy\u003C\u002Fli>\u003Cli>The reported gain is an average 1.26x over naive Q pretraining\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you build RL systems, this paper is a reminder that initialization strategy is part of the algorithm, not just a setup detail. A value function trained on the wrong distribution can look good offline and still fail to support the policy you actually want online.\u003C\u002Fp>\u003Cp>That is especially relevant in fine-tuning workflows, where the base policy is already good enough to deploy or adapt, and the goal is to improve it without destabilizing learning. In that setting, the value function should help the next stage of optimization, not just reflect the past.\u003C\u002Fp>\u003Cp>The practical takeaway is not “never pretrain Q-functions.” It is “be careful about what data and what policy distribution that pretraining represents.” If the offline target is too tightly coupled to the base policy, you may be paying for a mismatch.\u003C\u002Fp>\u003Ch2>Limitations and open questions\u003C\u002Fh2>\u003Cp>The paper’s own abstract is focused on continuous control benchmarks, so the evidence here is strongest for that setting. It does not claim, at least in the abstract, that IPE solves all value-based RL fine-tuning problems.\u003C\u002Fp>\u003Cp>We also do not get benchmark numbers beyond the average 1.26x improvement, so it is hard to judge variance, task sensitivity, or how much of the gain comes from specific environments. Those details would matter for teams deciding whether to adopt the method in production-like training loops.\u003C\u002Fp>\u003Cp>There is also a broader open question: how many diverse policies do you need, and how should you choose them? The abstract says “multiple diverse policies,” but not how diversity is measured or how sensitive the method is to that design choice.\u003C\u002Fp>\u003Cp>Still, the paper makes a useful point for RL engineers: if your fine-tuning pipeline assumes the offline Q-function should just mirror the pretrained policy, that assumption may be doing more harm than good. IPE is a lightweight alternative worth understanding if you work on online adaptation, control, or policy improvement loops.\u003C\u002Fp>","A new method says online RL fine-tuning can work better when Q-functions start from diverse policy rollouts.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.27203",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785391372788-de5m.png","research","en","08ceac3e-dd49-42e7-b976-e962eae021ca",[17,18,19,20,21],"reinforcement learning","Q-functions","fine-tuning","continuous control","policy ensemble",[23,24,25],"Naive Q-function pretraining often adds little over random initialization.","The mismatch comes from pretraining targeting the old policy, not the fine-tuned one.","IPE bootstraps Q-learning from pooled rollouts of multiple diverse policies.",0,"2026-07-30T06:02:24.317984+00:00","2026-07-30T06:02:24.311+00:00",{"tags":30,"relatedLang":34,"relatedPosts":38},[31,32],{"name":19,"slug":19},{"name":17,"slug":33},"reinforcement-learning",{"id":15,"slug":35,"title":36,"language":37},"pretrain-q-functions-online-rl-finetuning-zh","Q 函數不一定要先預訓練","zh",[39,45,51,57,63,69],{"id":40,"slug":41,"title":42,"cover_image":43,"image_url":43,"created_at":44,"category":13},"8e53848b-d163-4335-a8e4-29694e85bbb3","fruitfly-inspired-regression-without-heavy-models-en","Fruitfly-Inspired Regression Without Heavy Models","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785394974203-76p6.png","2026-07-30T07:02:27.641563+00:00",{"id":46,"slug":47,"title":48,"cover_image":49,"image_url":49,"created_at":50,"category":13},"2515b20a-f125-4354-b386-50e75eff70c4","mental-world-modeling-simulating-minds-en","Mental World Modeling: Simulating minds, not just scenes","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785393182740-cnwi.png","2026-07-30T06:32:29.590725+00:00",{"id":52,"slug":53,"title":54,"cover_image":55,"image_url":55,"created_at":56,"category":13},"654f2009-0838-4f91-946b-61e508f5ba9b","openai-agent-hack-forces-tighter-eval-controls-en","OpenAI’s agent hack forces tighter eval controls","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785326591602-45xj.png","2026-07-29T12:02:46.361703+00:00",{"id":58,"slug":59,"title":60,"cover_image":61,"image_url":61,"created_at":62,"category":13},"459e2d94-412c-472f-991b-1fe9d42bb684","care-confidence-adaptive-routing-lora-en","CARE routes LoRA experts by confidence","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785308570440-rdwq.png","2026-07-29T07:02:29.806618+00:00",{"id":64,"slug":65,"title":66,"cover_image":67,"image_url":67,"created_at":68,"category":13},"4347dd8c-0949-4bd2-9e36-bbbf1d467b4b","pir2-reactive-real-time-flow-policies-en","πR² makes flow policies react in real time","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785306784482-fogi.png","2026-07-29T06:32:34.74935+00:00",{"id":70,"slug":71,"title":72,"cover_image":73,"image_url":73,"created_at":74,"category":13},"e2f21eaf-1f13-4f5a-9796-87b499de7422","relay-opd-fixes-prefix-failure-distillation-en","Relay-OPD fixes prefix failure in distillation","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785304980702-n7zz.png","2026-07-29T06:02:31.313404+00:00",[76,81,86,91,96,101,106,111,116,121],{"id":77,"slug":78,"title":79,"created_at":80},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":82,"slug":83,"title":84,"created_at":85},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":87,"slug":88,"title":89,"created_at":90},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":92,"slug":93,"title":94,"created_at":95},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":97,"slug":98,"title":99,"created_at":100},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":102,"slug":103,"title":104,"created_at":105},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":107,"slug":108,"title":109,"created_at":110},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":112,"slug":113,"title":114,"created_at":115},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":117,"slug":118,"title":119,"created_at":120},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":122,"slug":123,"title":124,"created_at":125},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]