[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-fine-tune-small-llm-legal-labeling-en":3,"article-related-fine-tune-small-llm-legal-labeling-en":29,"series-ai-agent-445ecce7-ea09-49f7-a7f4-88391eb1bbf3":73},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"445ecce7-ea09-49f7-a7f4-88391eb1bbf3","fine-tune-small-llm-legal-labeling-en","Fine-tune a small LLM for legal labeling","\u003Cp data-speakable=\"summary\">81.7% shows a 3B SmolLM3 can beat frontier models on one legal-labeling task.\u003C\u002Fp>\u003Cp>This guide is for developers who want a practical path from a general open model to a specialist that handles one repeated decision well.\u003C\u002Fp>\u003Cp>By the end, you will have a fine-tuned small language model, a repeatable evaluation loop, and a clear rule for when to route hard requests to a larger model.\u003C\u002Fp>\u003Ch2>Before you start\u003C\u002Fh2>\u003Cul>\u003Cli>An account with access to a GPU host or local machine with an NVIDIA GPU.\u003C\u002Fli>\u003Cli>Python 3.10+ and pip 23+.\u003C\u002Fli>\u003Cli>Node is not required.\u003C\u002Fli>\u003Cli>Hugging Face account and access token.\u003C\u002Fli>\u003Cli>PyTorch 2.2+ with CUDA 12.1 or newer.\u003C\u002Fli>\u003Cli>Transformers 4.40+, Datasets 2.19+, PEFT 0.11+, and Accelerate 0.30+.\u003C\u002Fli>\u003Cli>A labeled legal contract dataset in CSV or JSONL form.\u003C\u002Fli>\u003Cli>Enough disk space for the base model, adapter weights, and evaluation set.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Step 1: Pick a narrow legal task\u003C\u002Fh2>\u003Cp>Your first outcome is a task definition that a small model can learn from examples instead of broad reasoning.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786150968314-7umc.png\" alt=\"Fine-tune a small LLM for legal labeling\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Choose one repeated decision, such as contract clause tagging, support-ticket routing, or legal issue classification. Keep the label set small and stable, and write one sentence that defines the input and one sentence that defines the expected output.\u003C\u002Fp>\u003Cp>For example, use contract text as input and one of a fixed set of tags as output. If humans cannot label the same example consistently, the model will not learn a clean pattern.\u003C\u002Fp>\u003Cp>You should see a task spec with clear labels, example inputs, and a short success criterion such as “match human tags on held-out contracts.”\u003C\u002Fp>\u003Ch2>Step 2: Prepare the dataset for LoRA fine-tuning\u003C\u002Fh2>\u003Cp>Your second outcome is a training file the model can consume without extra cleanup during training.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786150966826-gliw.png\" alt=\"Fine-tune a small LLM for legal labeling\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Split your data into train, validation, and test sets. Remove duplicates, normalize label names, and convert each record into a prompt-response pair. If you use JSONL, keep one example per line so you can stream the data into training.\u003C\u002Fp>\u003Cpre>\u003Ccode>{\"prompt\":\"Classify this clause: ...\",\"response\":\"termination\"}\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>Make sure the validation and test sets contain contracts the model has never seen. You should see counts for each split and no overlap between them.\u003C\u002Fp>\u003Ch2>Step 3: Fine-tune SmolLM3 with LoRA\u003C\u002Fh2>\u003Cp>Your third outcome is an adapter checkpoint that teaches the base model your legal labels without retraining every weight.\u003C\u002Fp>\u003Cp>Start from a small open base model such as SmolLM3, then attach LoRA adapters and train only the low-rank layers. This keeps the run cheap enough for a single \u003Ca href=\"\u002Ftag\u002Fgpu\">GPU\u003C\u002Fa> and makes it practical to iterate on the dataset.\u003C\u002Fp>\u003Cp>Use your framework of choice to launch training, then save only the adapter weights and tokenizer files. Keep the base model frozen so you can compare runs fairly.\u003C\u002Fp>\u003Cp>You should see training loss fall, validation loss stabilize, and an adapter folder appear at the end of the run.\u003C\u002Fp>\u003Ch2>Step 4: Measure against frontier baselines\u003C\u002Fh2>\u003Cp>Your fourth outcome is an honest scorecard that tells you whether the specialist model is actually useful.\u003C\u002Fp>\u003Cp>Run the same held-out test set through the fine-tuned model and through one or two frontier baselines, then compare exact-match accuracy or macro F1. In the source \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa>, the fine-tuned 3B model finished at 81.7% after 74 minutes, ahead of GPT-5.5 at 76.7% and \u003Ca href=\"\u002Ftag\u002Fclaude\">Claude\u003C\u002Fa> Sonnet 4.6 at 77% on that narrow task.\u003C\u002Fp>\u003Cp>Track both quality and operational cost. A small model can win on one workflow even if it is not better at open-ended reasoning. You should see a table of predictions, scores by model, and a clear winner for your target task.\u003C\u002Fp>\u003Ch2>Step 5: Add routing for hard cases\u003C\u002Fh2>\u003Cp>Your fifth outcome is a production path that keeps cheap requests on the small model and escalates edge cases to a larger one.\u003C\u002Fp>\u003Cp>Set a confidence threshold or a rule-based fallback. If the small model is uncertain, send the request to a frontier model and log both decisions. This gives you low latency for common cases and a safety net for ambiguous ones.\u003C\u002Fp>\u003Cp>Keep the router simple at first. You can route by score margin, label entropy, or a classifier trained on past failures. You should see most traffic handled by the small model and a smaller share forwarded upstream.\u003C\u002Fp>\u003Ch2>Step 6: Ship a monitoring loop\u003C\u002Fh2>\u003Cp>Your sixth outcome is a system that stays accurate after the first release.\u003C\u002Fp>\u003Cp>Create a weekly review set from real production examples, then re-run the same evaluation script after every dataset or prompt change. Watch for label drift, hallucinated tags, and regressions after adapter updates.\u003C\u002Fp>\u003Cp>If performance drops, retrain on fresh examples before expanding the model’s scope. You should see a stable dashboard with live accuracy, fallback rate, and a list of recent failures.\u003C\u002Fp>\u003Ctable>\u003Cthead>\u003Ctr>\u003Cth>Metric\u003C\u002Fth>\u003Cth>Before\u002FBaseline\u003C\u002Fth>\u003Cth>After\u002FResult\u003C\u002Fth>\u003C\u002Ftr>\u003C\u002Fthead>\u003Ctbody>\u003Ctr>\u003Ctd>Legal-labeling accuracy\u003C\u002Ftd>\u003Ctd>GPT-5.5: 76.7%\u003C\u002Ftd>\u003Ctd>SmolLM3 fine-tune: 81.7%\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Legal-labeling accuracy\u003C\u002Ftd>\u003Ctd>Claude Sonnet 4.6: 77%\u003C\u002Ftd>\u003Ctd>SmolLM3 fine-tune: 81.7%\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd>Training time\u003C\u002Ftd>\u003Ctd>Not reported\u003C\u002Ftd>\u003Ctd>74 minutes\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch2>Common mistakes\u003C\u002Fh2>\u003Cul>\u003Cli>Using a broad task with vague labels. Fix: narrow the scope to one repeated decision and define labels before training.\u003C\u002Fli>\u003Cli>Training on noisy or duplicated examples. Fix: deduplicate, normalize labels, and hold out a clean test set.\u003C\u002Fli>\u003Cli>Skipping fallback routing. Fix: send low-confidence requests to a frontier model and log every escalation.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>What's next\u003C\u002Fh2>\u003Cp>After this, try adding retrieval for fresh policy text, distillation for cheaper \u003Ca href=\"\u002Ftag\u002Finference\">inference\u003C\u002Fa>, or quantization so the same specialist can run on smaller hardware.\u003C\u002Fp>","A 3B SmolLM3 model reached 81.7% on legal labeling after 74 minutes of fine-tuning.","newsletter.systemdesign.one","https:\u002F\u002Fnewsletter.systemdesign.one\u002Fp\u002Ffine-tuning-small-language-models",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1786150968314-7umc.png","ai-agent","en","10d812f4-bf80-40c4-9572-3f2346b1234d",[17,18,19,20,21],"SmolLM3","LoRA","fine-tuning","legal labeling","model routing",[23,24,25],"A small model can outperform frontier models on one narrow workflow when it is trained on the right examples.","LoRA makes specialist fine-tuning practical on a single GPU by updating only a small set of parameters.","Production use needs routing and monitoring so hard cases can fall back to a larger model.",1,"2026-08-08T01:02:27.269166+00:00","2026-08-08T01:02:27.267+00:00",{"tags":30,"relatedLang":32,"relatedPosts":36},[31],{"name":19,"slug":19},{"id":15,"slug":33,"title":34,"language":35},"fine-tune-small-llm-legal-labeling-zh","法律標註微調小型 LLM 產出","zh",[37,43,49,55,61,67],{"id":38,"slug":39,"title":40,"cover_image":41,"image_url":41,"created_at":42,"category":13},"306b4d13-3911-4fcb-9f53-861fa5e9b430","sala-boosts-long-context-edge-ai-en","SALA Boosts Long Context on Edge AI","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785765793459-52io.png","2026-08-03T14:02:45.076761+00:00",{"id":44,"slug":45,"title":46,"cover_image":47,"image_url":47,"created_at":48,"category":13},"33146ff0-fa4b-4174-9d11-e0dcbf120bdc","anthropic-breach-proves-ai-agents-need-hard-security-limits-en","Anthropic’s breach proves AI agents need hard security limits","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785742391184-iku1.png","2026-08-03T07:32:42.32307+00:00",{"id":50,"slug":51,"title":52,"cover_image":53,"image_url":53,"created_at":54,"category":13},"0a805e08-8c91-43a5-b6e2-a028c60c27a8","genai-mil-war-prompt-report-template-en","GenAI.mil turns a scary prompt into a report","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785655981099-6hcj.png","2026-08-02T07:32:38.593628+00:00",{"id":56,"slug":57,"title":58,"cover_image":59,"image_url":59,"created_at":60,"category":13},"a7767079-eebd-4221-a26a-3d55dacfdb5c","epam-openai-deal-turns-pilots-into-production-en","EPAM’s OpenAI deal turns pilots into production","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785628985506-duom.png","2026-08-02T00:02:40.569745+00:00",{"id":62,"slug":63,"title":64,"cover_image":65,"image_url":65,"created_at":66,"category":13},"05a6dd16-f95c-464a-9c59-14c1c3c67487","prompt-engineering-overrated-claude-code-en","Prompt engineering is overrated for Claude Code","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785544368755-npgt.png","2026-08-01T00:32:24.626964+00:00",{"id":68,"slug":69,"title":70,"cover_image":71,"image_url":71,"created_at":72,"category":13},"cc48d966-f3b8-42b9-aca8-0accf83394bc","grok-build-live-previews-rewind-fixes-en","Grok Build adds live previews and rewind fixes","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785198775225-m3fo.png","2026-07-28T00:32:31.401047+00:00",[74,79,84,89,94,99,104,109,114,119],{"id":75,"slug":76,"title":77,"created_at":78},"03db8de8-8dc2-4ac1-9cf7-898782efbb1f","anthropic-claude-ai-agent-task-automation-en","Anthropic's Claude AI Agent: A New Era of Task Automation","2026-03-25T16:25:06.513026+00:00",{"id":80,"slug":81,"title":82,"created_at":83},"045d1abc-190d-4594-8c95-91e2a26f0c5a","googles-2026-ai-agent-report-decoded-en","Google’s 2026 AI Agent Report, Decoded","2026-03-26T11:15:23.046616+00:00",{"id":85,"slug":86,"title":87,"created_at":88},"e64aba21-254b-4f93-aa21-837484bb52ec","kimi-k25-review-stronger-still-not-legend-en","Kimi K2.5 review: stronger, still not a legend","2026-03-27T07:15:55.385951+00:00",{"id":90,"slug":91,"title":92,"created_at":93},"30dfb781-a1b2-4add-aebe-b3df40247c37","claude-code-controls-mac-desktop-en","Claude Code now controls your Mac desktop","2026-03-28T03:01:59.384091+00:00",{"id":95,"slug":96,"title":97,"created_at":98},"254405b6-7833-4800-8e13-f5196deefbe6","cloudflare-100x-faster-ai-agent-sandbox-en","Cloudflare’s 100x Faster AI Agent Sandbox","2026-03-28T03:09:44.356437+00:00",{"id":100,"slug":101,"title":102,"created_at":103},"04f29b7f-9b91-4306-89a7-97d725e6e1ba","openai-backs-isara-agent-swarm-bet-en","OpenAI backs Isara’s agent-swarm bet","2026-03-28T03:15:27.849766+00:00",{"id":105,"slug":106,"title":107,"created_at":108},"3b0bf479-e4ae-4703-9666-721a7e0cdb91","openai-plan-automated-ai-researcher-en","OpenAI’s plan for an automated AI researcher","2026-03-28T03:17:42.312819+00:00",{"id":110,"slug":111,"title":112,"created_at":113},"fe91bce0-b85d-4efa-a207-24ae9939c29f","harness-engineering-ai-agent-reliability-2026","Harness Engineering: From Bridle to Operating System, The Missing Link in AI Agent Reliability","2026-03-31T06:36:55.648751+00:00",{"id":115,"slug":116,"title":117,"created_at":118},"7a09007d-820f-43b3-8607-8ad1bfcb94c8","mcp-explained-from-prompts-to-production-en","MCP Explained: From Prompts to Production","2026-04-01T09:24:40.089177+00:00",{"id":120,"slug":121,"title":122,"created_at":123},"116d5ee9-a4f1-4b5a-aac5-5d035dd22bbe","amazon-bedrock-agents-multi-agent-workflows-en","Amazon Bedrock Agents Gets Multi-Agent Workflows","2026-04-01T09:30:30.197685+00:00"]