[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-coinrag-fine-grained-kv-cache-reuse-rag-en":3,"article-related-coinrag-fine-grained-kv-cache-reuse-rag-en":29,"series-research-1f0b474d-49e9-4ce1-a3bb-a6b6561ed107":78},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"1f0b474d-49e9-4ce1-a3bb-a6b6561ed107","coinrag-fine-grained-kv-cache-reuse-rag-en","CoinRAG Reuses Fine-Grained KV Caches for RAG","\u003Cp data-speakable=\"summary\">CoinRAG cuts long-context \u003Ca href=\"\u002Ftag\u002Frag\">RAG\u003C\u002Fa> prefill cost by reusing fine-grained nugget KV caches instead of full chunks.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 5.3% relative improvement in answer quality (F1)\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Two-stage retrieval assembles sliced nugget KV with chunk-level context\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Long-context RAG systems often pay a steep price just to read retrieved passages before they can answer a question. CoinRAG is trying to make that first pass cheaper without throwing away the semantic signal that actually matters.\u003C\u002Fp>\u003Cp>The practical angle is simple: if a retriever brings back a lot of text, not all of it deserves equal compute. This paper argues that chunk-level cache reuse is still too coarse, because chunks can contain redundancy and noise that waste prefill budget while doing little for accuracy.\u003C\u002Fp>\u003Ch2>What problem CoinRAG is solving\u003C\u002Fh2>\u003Cp>Retrieval-augmented generation usually works by pulling in external context and then feeding that context into the model. For long retrieved passages, the prefill stage can become expensive, especially when the system has to process many tokens before it can start generating an answer.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786345380961-irlf.png\" alt=\"CoinRAG Reuses Fine-Grained KV Caches for RAG\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Earlier optimization work tried to reduce that cost by reusing \u003Ca href=\"\u002Ftag\u002Fkv-cache\">KV cache\u003C\u002Fa> at the chunk level. That helps, but the abstract says coarse chunks still leave a lot of redundant or irrelevant material in the path. In other words, the system may be saving compute, but it is still spending it on text that does not pull its weight.\u003C\u002Fp>\u003Cp>CoinRAG targets that gap by trying to optimize the Pareto frontier under low prefill latency constraints while preserving accuracy. For engineers, that means the paper is about a familiar systems tradeoff: getting faster without collapsing answer quality.\u003C\u002Fp>\u003Ch2>How the method works in plain English\u003C\u002Fh2>\u003Cp>The core idea is to move from whole chunks to smaller semantic units the paper calls “information nuggets.” Instead of encoding an entire retrieved chunk in one shot, CoinRAG identifies query-relevant nuggets inside the chunk and reuses offline-computed KV caches for those smaller pieces.\u003C\u002Fp>\u003Cp>The abstract describes this as a compositional process. The system does not just drop in isolated nugget caches; it assembles sliced KV representations together with chunk-level context so the model still sees a coherent representation rather than a pile of disconnected fragments.\u003C\u002Fp>\u003Cp>That is the key design choice here: keep the contextual benefits of retrieval, but spend compute only on the parts of the chunk that matter most to the current query. The method is framed as a learned contextual representation that is more semantically relevant and more compact than full-chunk encoding.\u003C\u002Fp>\u003Cp>There is also a two-stage retrieval step. The abstract says CoinRAG first identifies query-relevant semantic units within retrieved chunks, then combines their sliced KV representations with the broader chunk context. The details of how those stages are implemented are not spelled out in the abstract, so the exact selection and assembly logic would need the full paper.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The evaluation is on LongBench multi-hop question answering tasks. The abstract does not list the full \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> suite, per-task breakdowns, or latency numbers, so those details are not available from the source material here.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786345386030-ngc6.png\" alt=\"CoinRAG Reuses Fine-Grained KV Caches for RAG\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>What it does claim is that CoinRAG significantly reduces operational costs and outperforms the baselines, producing a new Pareto frontier under a standard fast prefill latency budget. The single concrete metric given is an average 5.3% relative improvement in answer quality measured by F1.\u003C\u002Fp>\u003Cp>That matters because the result is not just “faster” or just “more accurate.” The paper is claiming a better balance of both under a latency constraint. In systems terms, that is often the difference between a neat optimization and something you can actually deploy.\u003C\u002Fp>\u003Cp>Still, the abstract leaves out important context. We do not get the absolute latency savings, memory savings, or the exact baselines beaten. We also do not see whether the gains are uniform across all query types or concentrated in certain multi-hop cases.\u003C\u002Fp>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you are building RAG systems over large retrieved contexts, the cost center is often prefill, not generation. Any method that trims prefill without wrecking answer quality is directly relevant to throughput, latency, and serving cost.\u003C\u002Fp>\u003Cp>CoinRAG is especially interesting because it attacks the granularity problem. Chunk-level reuse is easy to reason about, but it can be wasteful when only a few spans inside a chunk are actually useful. Fine-grained KV reuse is a more surgical approach, and this paper suggests that the extra complexity may pay off.\u003C\u002Fp>\u003Cp>For practitioners, the likely takeaway is architectural rather than plug-and-play: retrieval pipelines may benefit from treating context as a set of reusable semantic units instead of a monolithic blob. That could influence how you design indexing, caching, and prompt assembly for long-context workloads.\u003C\u002Fp>\u003Ch2>Limitations and open questions\u003C\u002Fh2>\u003Cp>The abstract is strong on the high-level idea but light on implementation details. It does not explain the nugget extraction algorithm, the offline cache construction cost, or how much extra engineering is required to maintain those sliced KV caches.\u003C\u002Fp>\u003Cp>It also does not tell us how robust the method is outside LongBench multi-hop QA. So while the paper reports a better Pareto frontier on that benchmark family, it is still an open question how general the gains are across other RAG settings, model sizes, or retrieval qualities.\u003C\u002Fp>\u003Cp>Another practical question is operational overhead. Fine-grained cache reuse sounds efficient at \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa> time, but it can introduce more bookkeeping during preprocessing and retrieval orchestration. The abstract does not provide enough detail to judge that tradeoff.\u003C\u002Fp>\u003Cp>Even with those gaps, the direction is clear: CoinRAG is pushing RAG optimization from coarse chunk reuse toward semantic, reusable subchunk representations. For teams trying to squeeze more quality out of fixed latency budgets, that is exactly the kind of idea worth watching.\u003C\u002Fp>\u003Cul>\u003Cli>CoinRAG targets the prefill bottleneck in long-context RAG.\u003C\u002Fli>\u003Cli>It reuses fine-grained nugget KV caches instead of full chunks.\u003C\u002Fli>\u003Cli>It reports a 5.3% average F1 gain under a fast prefill budget.\u003C\u002Fli>\u003C\u002Ful>","CoinRAG cuts long-context RAG prefill cost by reusing fine-grained nugget KV caches instead of full chunks.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.07458",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786345380961-irlf.png","research","en","01ff45d6-76b3-4cdb-99bf-95d620b383fb",[17,18,19,20,21],"RAG","KV cache","long-context","prefill latency","retrieval",[23,24,25],"Fine-grained nugget reuse is more selective than chunk-level cache reuse.","The paper reports a new Pareto frontier on LongBench multi-hop QA.","The abstract does not provide absolute latency or memory numbers.",1,"2026-08-10T07:02:34.138838+00:00","2026-08-10T07:02:34.136+00:00",{"tags":30,"relatedLang":37,"relatedPosts":41},[31,33,35],{"name":17,"slug":32},"rag",{"name":18,"slug":34},"kv-cache",{"name":36,"slug":19},"long context",{"id":15,"slug":38,"title":39,"language":40},"coinrag-fine-grained-kv-cache-reuse-rag-zh","CoinRAG 用細粒度 KV 快取加速 RAG","zh",[42,48,54,60,66,72],{"id":43,"slug":44,"title":45,"cover_image":46,"image_url":46,"created_at":47,"category":13},"d26d3c47-b7f3-4598-a877-e7be8c67cf67","creativeinstruct-llms-quality-creativity-diversity-en","CreativeInstruct teaches LLMs to stay creative","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786343571807-rqw1.png","2026-08-10T06:32:28.086694+00:00",{"id":49,"slug":50,"title":51,"cover_image":52,"image_url":52,"created_at":53,"category":13},"30ce677e-59ab-4f83-9759-da93aa8bb4af","mirrorworld-mirror-reflection-video-diffusion-en","MirrorWorld makes mirror reflections consistent in video","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786341775274-6d5a.png","2026-08-10T06:02:27.25669+00:00",{"id":55,"slug":56,"title":57,"cover_image":58,"image_url":58,"created_at":59,"category":13},"3f886925-6381-4770-980a-1001203cf245","claude-4-5-ai-progress-still-accelerating-en","Claude 4.5 proves AI progress is still accelerating","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786257168594-qg8m.png","2026-08-09T06:32:27.560501+00:00",{"id":61,"slug":62,"title":63,"cover_image":64,"image_url":64,"created_at":65,"category":13},"8675701f-d283-4d69-8818-bdf59e5ae09e","mage-vl-cuts-visual-tokens-by-reading-codecs-en","Mage-VL Cuts Visual Tokens by Reading Codecs","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786170812146-kqtv.png","2026-08-08T06:33:06.377112+00:00",{"id":67,"slug":68,"title":69,"cover_image":70,"image_url":70,"created_at":71,"category":13},"7368755d-86ca-461e-9d95-d7e74e95b561","astra-turns-long-math-tasks-into-multi-agent-work-en","Astra turns long math tasks into multi-agent work","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786088006851-29p0.png","2026-08-07T07:32:52.200646+00:00",{"id":73,"slug":74,"title":75,"cover_image":76,"image_url":76,"created_at":77,"category":13},"e69199db-e1f8-4e12-aaf2-ea92eeb2e0cc","evidence-linked-feature-engineering-heart-failure-en","Evidence-linked feature engineering for heart failure","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786086190016-bykl.png","2026-08-07T07:02:31.382531+00:00",[79,84,89,94,99,104,109,114,119,124],{"id":80,"slug":81,"title":82,"created_at":83},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":85,"slug":86,"title":87,"created_at":88},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":90,"slug":91,"title":92,"created_at":93},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":95,"slug":96,"title":97,"created_at":98},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":100,"slug":101,"title":102,"created_at":103},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":105,"slug":106,"title":107,"created_at":108},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":110,"slug":111,"title":112,"created_at":113},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":115,"slug":116,"title":117,"created_at":118},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":120,"slug":121,"title":122,"created_at":123},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":125,"slug":126,"title":127,"created_at":128},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]