[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-kimi-k3-gpu-cost-self-hosted-vs-api-en":3,"article-related-kimi-k3-gpu-cost-self-hosted-vs-api-en":30,"series-industry-e113cf9f-096b-4742-b7bf-1c5b4b18ec3f":73},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":23,"views":27,"created_at":28,"published_at":29,"topic_cluster_id":11},"e113cf9f-096b-4742-b7bf-1c5b4b18ec3f","kimi-k3-gpu-cost-self-hosted-vs-api-en","Kimi K3 Needs About 1.5 TB of Memory","\u003Cp data-speakable=\"summary\">Kimi K3 needs about 1.5 TB of memory, so most teams will find \u003Ca href=\"\u002Fnews\u002Fopenai-api-pricing-august-2026-token-costs-en\">API pricing\u003C\u002Fa> far cheaper than self-hosted GPU clusters.\u003C\u002Fp>\u003Cp>Kimi K3 is a \u003Ca href=\"https:\u002F\u002Fmoonshot.ai\" target=\"_blank\" rel=\"noopener\">Moonshot AI\u003C\u002Fa> model with 2.8 trillion total parameters, but its deployment math is shaped by much more than raw size. The article’s core claim is simple: if you want to run Kimi K3 yourself, you are planning around roughly 1.5 TB of memory, multiple high-end GPUs, and a very specific software stack.\u003C\u002Fp>\u003Cp>That makes the real decision less about “can it run?” and more about “when does self-hosting beat \u003Ca href=\"\u002Ftag\u002Fapi\">API\u003C\u002Fa> pricing?” The answer depends on \u003Ca href=\"\u002Ftag\u002Ftoken\">token\u003C\u002Fa> volume, concurrency, context length, and whether you need direct control over data and infrastructure.\u003C\u002Fp>\u003Ctable>\u003Cthead>\u003Ctr>\u003Cth>Metric\u003C\u002Fth>\u003Cth>Value\u003C\u002Fth>\u003Cth>Why it matters\u003C\u002Fth>\u003C\u002Ftr>\u003C\u002Fthead>\u003Ctbody>\u003Ctr>\u003Ctd>Total parameters\u003C\u002Ftd>\u003Ctd>2.8 trillion\u003C\u002Ftd>\u003Ctd>Explains why weight storage is enormous\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Active parameters per token\u003C\u002Ftd>\u003Ctd>About 104 billion\u003C\u002Ftd>\u003Ctd>Shows the MoE design keeps inference sparse\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>FP4 weight memory\u003C\u002Ftd>\u003Ctd>About 1.4 TB\u003C\u002Ftd>\u003Ctd>Just the model weights already exceed a single GPU\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Total memory per request\u003C\u002Ftd>\u003Ctd>About 1.5 TB\u003C\u002Ftd>\u003Ctd>Includes weights, KV cache, activations, and runtime overhead\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>API output price\u003C\u002Ftd>\u003Ctd>$15 per million tokens\u003C\u002Ftd>\u003Ctd>Sets the benchmark for cost comparisons\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch2>Why Kimi K3 is hard to host locally\u003C\u002Fh2>\u003Cp>Kimi K3 uses a Mixture of Experts design, so each token activates only a slice of the model. The article says each token uses 16 of 896 experts plus two shared experts, which keeps \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa> cheaper than a dense model of similar size.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786323794535-qyl7.png\" alt=\"Kimi K3 Needs About 1.5 TB of Memory\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>But sparse compute does not erase the storage problem. Kimi K3 ships with native MXFP4 weights, and the article estimates memory with a simple formula: total parameters multiplied by quantization bits, then divided by 8. On that basis, 2.8 trillion parameters at 4 bits comes out to about 1.4 TB for weights alone.\u003C\u002Fp>\u003Cp>That number gets bigger once you add the rest of the inference stack. The article adds another 2 to 15 GB for \u003Ca href=\"\u002Ftag\u002Fkv-cache\">KV cache\u003C\u002Fa>, around 30 GB for activations, and about 30 GB for runtime overhead. The total lands near 1.5 TB for a single request path.\u003C\u002Fp>\u003Cul>\u003Cli>Native FP4 weights reduce storage pressure compared with BF16 or FP16.\u003C\u002Fli>\u003Cli>Higher precision modes exist, but they are mainly for compatibility, not better output quality.\u003C\u002Fli>\u003Cli>Fine-tuning still needs higher precision, so FP4 is mostly an inference format.\u003C\u002Fli>\u003Cli>Kimi K2, by comparison, needs about 2 TB in BF16 for 1 trillion parameters.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>The GPU count is only half the story\u003C\u002Fh2>\u003Cp>Once the memory math is clear, the hardware question becomes blunt: a single GPU cannot host Kimi K3. The article estimates that you would need about 19 \u003Ca href=\"https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fdata-center\u002Fh100\u002F\" target=\"_blank\" rel=\"noopener\">NVIDIA H100\u003C\u002Fa> cards at 80 GB each, or about 11 \u003Ca href=\"https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fdata-center\u002Fh200\u002F\" target=\"_blank\" rel=\"noopener\">NVIDIA H200\u003C\u002Fa> cards at 141 GB each, just to cover the base memory requirement.\u003C\u002Fp>\u003Cp>That still does not mean those cards are the best fit. Hopper-class hardware and AMD MI300X can load 4-bit weights, but the article notes that they usually need to dequantize at runtime. That preserves some memory savings, yet it gives up part of the speed advantage built into Kimi K3’s native 4-bit design.\u003C\u002Fp>\u003Cblockquote>“The best way to run Kimi K3 is with 8 NVIDIA B300 or 8 AMD MI355X GPUs and tensor parallelism set to 8,” according to the \u003Ca href=\"https:\u002F\u002Fdocs.vllm.ai\" target=\"_blank\" rel=\"noopener\">vLLM\u003C\u002Fa> deployment guidance referenced in the article.\u003C\u002Fblockquote>\u003Cp>That recommendation says a lot about where the model is headed in practice. The article argues that self-hosting Kimi K3 is really a hardware-generation decision, not just a capacity decision. If your stack cannot run native MXFP4 well, you are already paying an efficiency penalty.\u003C\u002Fp>\u003Cp>There is also a software wrinkle. Kimi K3 uses a custom prefix-caching approach tied to its attention design, so standard \u003Ca href=\"\u002Ftag\u002Fvllm\">vLLM\u003C\u002Fa> prefix caching does not fit cleanly. Moonshot contributed a custom implementation, which means self-hosters need a newer vLLM build with K3-specific support.\u003C\u002Fp>\u003Ch2>Where the cost curve flips\u003C\u002Fh2>\u003Cp>The article’s pricing comparison is the part most teams will care about. Kimi K3’s official API pricing is $3 per million input tokens, $0.30 per million cached input tokens, and $15 per million output tokens. \u003Ca href=\"https:\u002F\u002Fwww.digitalocean.com\u002Fproducts\u002Fserverless-inference\" target=\"_blank\" rel=\"noopener\">DigitalOcean Serverless Inference\u003C\u002Fa> uses the same input and output prices, so the main difference is whether you pay per token or keep paying for GPUs all month.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786323790275-xq35.png\" alt=\"Kimi K3 Needs About 1.5 TB of Memory\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>For a native FP4 self-hosted setup, the article uses an 8-GPU \u003Ca href=\"https:\u002F\u002Fwww.digitalocean.com\u002Fproducts\u002Fgpu-droplets\" target=\"_blank\" rel=\"noopener\">GPU Droplet\u003C\u002Fa> node built from AMD MI350X cards. At about $4.76 per GPU hour, the node costs roughly $38 per hour, or about $27,800 per month if it runs continuously.\u003C\u002Fp>\u003Cp>That monthly bill is hard to justify unless your workload is very large and very steady. The article estimates that to match that spend through API output pricing alone, you would need about 1.8 billion output tokens per month, which works out to around 700 tokens per second every hour of the month.\u003C\u002Fp>\u003Cul>\u003Cli>One heavy user at 50 million tokens per month would pay about $750 through the API if all tokens were billed as output.\u003C\u002Fli>\u003Cli>The same user would be far below the utilization needed to justify a dedicated 8-GPU node.\u003C\u002Fli>\u003Cli>The article says API use is about 40 times cheaper for a single heavy user under its assumptions.\u003C\u002Fli>\u003Cli>Self-hosting starts to make more sense once you have more than 40 steady heavy users.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That last point matters because “40 users” is really shorthand for sustained throughput. It could be 40 people, or it could be one engineer running a swarm of agents that keep the node busy all day.\u003C\u002Fp>\u003Ch2>Concurrency changes the math, but not enough for one user\u003C\u002Fh2>\u003Cp>The article also explains why a single user cannot simply rent fewer GPUs and call it a day. Tensor parallelism works best on 2, 4, or 8 GPUs, not odd counts like 6, and most providers sell these machines in fixed blocks anyway.\u003C\u002Fp>\u003Cp>Even if you could assemble a smaller set, token output rates would still be the bottleneck. A single request rarely produces enough tokens per second to keep a 6-GPU setup busy enough to offset its rental cost.\u003C\u002Fp>\u003Cp>What does help is multi-request concurrency. The article says one node can handle about 40 parallel requests, and with shorter contexts the number of resident requests can rise sharply because KV cache usage drops.\u003C\u002Fp>\u003Cul>\u003Cli>Full 1 million token context: about 15 GB per request, around 40 requests per 600 GB pool.\u003C\u002Fli>\u003Cli>128K context: about 2 GB per request, around 300 requests.\u003C\u002Fli>\u003Cli>32K context: about 1.5 GB per request, around 400 requests.\u003C\u002Fli>\u003Cli>8K context: about 1 GB per request, around 600 requests.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That is the real tradeoff. Kimi K3 becomes more attractive when you can keep the node busy with many requests, many agents, or long-lived enterprise workloads. If you only have one or two users, the API is still the cleaner financial choice.\u003C\u002Fp>\u003Ch2>What this means for teams choosing a deployment model\u003C\u002Fh2>\u003Cp>There is a practical split here. If you want the simplest path, token-based API access wins. You avoid buying expensive GPUs, you skip model serving work, and you do not need to maintain a custom inference stack.\u003C\u002Fp>\u003Cp>If you need tighter control over data residency, latency consistency, or model ownership, self-hosting starts to look better. The article also points out that some third-party providers do not store inference data, so privacy concerns alone do not automatically force a self-hosted setup.\u003C\u002Fp>\u003Cp>Licensing matters too. Kimi K3 does not use Apache or MIT licensing. It uses a custom license that allows use, modification, fine-tuning, deployment, distribution, and commercialization, but it adds conditions for large-scale model-as-a-service businesses and very large commercial products.\u003C\u002Fp>\u003Cp>For teams evaluating \u003Ca href=\"\u002Fnews\u002Fkimi-k3-deployment-costs\" target=\"_blank\" rel=\"noopener\">Kimi K3 deployment costs\u003C\u002Fa>, the takeaway is straightforward: compare token volume, not just token price. If your workload is bursty, the API is probably cheaper. If your workload is steady, multi-user, and GPU-efficient, dedicated infrastructure can eventually win.\u003C\u002Fp>\u003Cp>The more interesting question is what happens when more open-weight models arrive with similar FP4 support. If vendors keep shipping native low-precision inference paths, the gap between “cheap to call” and “cheap to own” will keep narrowing, but the break-even point will still be determined by utilization, not hype.\u003C\u002Fp>","Kimi K3 needs about 1.5 TB of memory, so most teams will find API pricing far cheaper than self-hosted GPU clusters.","zhuanlan.zhihu.com","https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F2067640378094785740",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786323794535-qyl7.png","industry","en","fa5bffe8-c724-42ba-bbe5-94591fb9e766",[17,18,19,20,21,22],"Kimi K3","GPU deployment","API pricing","self-hosting","MXFP4","vLLM",[24,25,26],"Kimi K3 needs about 1.5 TB of memory for a practical self-hosted inference setup.","A single heavy user is usually far cheaper on API billing than on rented GPUs.","Self-hosting becomes interesting only when usage is steady, parallel, and large enough to keep 8-GPU nodes busy.",2,"2026-08-10T01:02:51.584977+00:00","2026-08-10T01:02:51.579+00:00",{"tags":31,"relatedLang":32,"relatedPosts":36},[],{"id":15,"slug":33,"title":34,"language":35},"kimi-k3-gpu-api-cost-comparison-zh","Kimi K3 自托管成本很高","zh",[37,43,49,55,61,67],{"id":38,"slug":39,"title":40,"cover_image":41,"image_url":41,"created_at":42,"category":13},"c714a29c-a09c-40bd-a0b4-59dcd4e80878","crypto-infrastructure-era-ai-agents-en","Crypto’s infrastructure era arrives with AI agents","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786347180651-mx60.png","2026-08-10T07:32:32.346357+00:00",{"id":44,"slug":45,"title":46,"cover_image":47,"image_url":47,"created_at":48,"category":13},"a5fe53a8-a51c-4005-b927-0bcd383c234c","ai-weekly-2026-w33-en","AI Weekly: 2026-08-03 ~ 2026-08-10","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786335625470-q29j.png","2026-08-10T04:00:28.478721+00:00",{"id":50,"slug":51,"title":52,"cover_image":53,"image_url":53,"created_at":54,"category":13},"e2aa152b-ca7a-4c8d-b477-9572fab7bce9","cursor-should-not-copy-vs-code-design-en","Cursor should not copy VS Code’s new design","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786278772811-6w0u.png","2026-08-09T12:32:20.905169+00:00",{"id":56,"slug":57,"title":58,"cover_image":59,"image_url":59,"created_at":60,"category":13},"84784a2f-d4ea-46ab-a000-f797e00f77a7","institutional-crypto-tops-30b-tokenization-en","Institutional Crypto Tops $30B in Tokenization","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786233767183-e9ni.png","2026-08-09T00:02:29.157244+00:00",{"id":62,"slug":63,"title":64,"cover_image":65,"image_url":65,"created_at":66,"category":13},"21288e9b-3748-49ef-80ed-2e1da7534e90","icons-claude-deal-shows-clinical-trials-need-ai-not-pilots-en","Icon’s Claude deal shows clinical trials need AI, not pilots","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786212159683-a38e.png","2026-08-08T18:02:16.717855+00:00",{"id":68,"slug":69,"title":70,"cover_image":71,"image_url":71,"created_at":72,"category":13},"3f78457f-ecfb-4796-a3e3-e58bcf816dfb","nvidia-ssi-deal-turns-compute-into-toll-booth-en","NVIDIA’s SSI deal turns compute into a toll booth","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786190594828-xkcj.png","2026-08-08T12:02:49.628514+00:00",[74,79,84,89,94,99,104,109,114,119],{"id":75,"slug":76,"title":77,"created_at":78},"d35a1bd9-e709-412e-a2df-392df1dc572a","ai-impact-2026-developments-market-en","AI's Impact in 2026: Key Developments and Market Shifts","2026-03-25T16:20:33.205823+00:00",{"id":80,"slug":81,"title":82,"created_at":83},"5ed27921-5fd6-492e-8c59-78393bf37710","trumps-ai-legislative-framework-en","Trump's AI Legislative Framework: What's Inside?","2026-03-25T16:22:20.005325+00:00",{"id":85,"slug":86,"title":87,"created_at":88},"e454a642-f03c-4794-b185-5f651aebbaca","nvidia-gtc-2026-key-highlights-innovations-en","NVIDIA GTC 2026: Key Highlights and Innovations","2026-03-25T16:22:47.882615+00:00",{"id":90,"slug":91,"title":92,"created_at":93},"0ebb5b16-774a-4922-945d-5f2ce1df5a6d","claude-usage-diversifies-learning-curves-en","Claude Usage Diversifies, Learning Curves Emerge","2026-03-25T16:25:50.770376+00:00",{"id":95,"slug":96,"title":97,"created_at":98},"69934e86-2fc5-4280-8223-7b917a48ace8","openclaw-ai-commoditization-concerns-en","OpenClaw's Rise Raises Concerns of AI Model Commoditization","2026-03-25T16:26:30.582047+00:00",{"id":100,"slug":101,"title":102,"created_at":103},"b4b2575b-2ac8-46b2-b90e-ab1d7c060797","google-gemini-ai-rollout-2026-en","Google's Gemini AI Rollout Extended to 2026","2026-03-25T16:28:14.808842+00:00",{"id":105,"slug":106,"title":107,"created_at":108},"6e18bc65-42ae-4ad0-b564-67d7f66b979e","meta-llama4-fabricated-results-scandal-en","Meta's Llama 4 Scandal: Fabricated AI Test Results Unveiled","2026-03-25T16:29:15.482836+00:00",{"id":110,"slug":111,"title":112,"created_at":113},"bf888e9d-08be-4f47-996c-7b24b5ab3500","accenture-mistral-ai-deployment-en","Accenture and Mistral AI Team Up for AI Deployment","2026-03-25T16:31:01.894655+00:00",{"id":115,"slug":116,"title":117,"created_at":118},"5382b536-fad2-49c6-ac85-9eb2bae49f35","mistral-ai-high-stakes-2026-en","Mistral AI: Facing High Stakes in 2026","2026-03-25T16:31:39.941974+00:00",{"id":120,"slug":121,"title":122,"created_at":123},"9da3d2d6-b669-4971-ba1d-17fdb3548ed5","cursors-meteoric-rise-pressures-en","Cursor's Meteoric Rise Faces Industry Pressures","2026-03-25T16:32:21.899217+00:00"]