{"id":145,"date":"2026-08-27T11:01:15","date_gmt":"2026-08-27T11:01:15","guid":{"rendered":"https:\/\/www.interviewbit.com\/varsity\/blog\/?p=145"},"modified":"2026-08-27T11:01:16","modified_gmt":"2026-08-27T11:01:16","slug":"fine-tuning-llms","status":"publish","type":"post","link":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/","title":{"rendered":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Fine-tuning an LLM means updating a pretrained model&#8217;s weights on your own examples so the new behaviour is built into the model instead of described in a prompt. In 2026 the practical stack is three layers: supervised fine-tuning (SFT) teaches the model a task and an output format from labelled examples, parameter-efficient fine-tuning (PEFT) methods such as LoRA and QLoRA do that by training a small set of added weights instead of all of them, and preference alignment methods such as RLHF and DPO tune which of several valid answers the model prefers. The rule that decides everything else: fine-tuning changes behaviour, retrieval changes knowledge. If your model is wrong about facts, you need RAG, not a fine-tune.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Most engineers arrive at fine-tuning after a prompt stops working, assume it is the next lever, and discover three weeks later that they spent GPU budget teaching a model to be confidently wrong in a new format. The technique is not hard. Knowing whether you need it is the hard part, and that is where this guide starts.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Fine-tuning is for form, not facts. Behaviour, format and style go in the weights; changing knowledge belongs in retrieval.<\/li>\n\n\n\n<li>The correct order is prompt, then RAG, then fine-tune. Skipping straight to training is the most common and most expensive mistake.<\/li>\n\n\n\n<li>LoRA and QLoRA are the only fine-tuning most teams should do. Full fine-tuning is rarely justified outside labs.<\/li>\n\n\n\n<li>Build the evaluation set before you train. Without it you cannot tell whether a checkpoint improved anything.<\/li>\n\n\n\n<li>RLHF is powerful and operationally heavy. DPO achieves preference alignment with far less machinery and is the sane default.<\/li>\n\n\n\n<li>The strongest commercial case for fine-tuning in 2026 is distillation: making a small cheap model match a frontier model on one narrow task.<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>First, Should You Fine-Tune at All?<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most teams that ask this question should not, at least not yet. The honest sequence is a ladder, and you climb it in order.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 1: Better prompting.<\/strong> A system message, a few examples, and an explicit output schema. This is free, instant, and reversible. It resolves more cases than engineers expect.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 2: Retrieval.<\/strong> If the model needs information it was not trained on, or information that changes, retrieval is the answer. Fine-tuning bakes a snapshot into the weights and goes stale the moment your data updates, and a fine-tuned model cannot cite the document that justified its answer.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 3: Tools.<\/strong> If the failure is arithmetic, lookups, or actions, give the model a tool instead of training it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 4: Fine-tune.<\/strong> Only now, and only when steps 1 to 3 have plateaued against a real evaluation set.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Three signals mean fine-tuning is genuinely the right call:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Format reliability.<\/strong> You need strict JSON or a fixed structure, and prompting gets you to 95% when you need 99.5%. Training collapses the variance that prompting cannot.<\/li>\n\n\n\n<li><strong>Cost and latency.<\/strong> A frontier model handles your task well but you call it a million times a month. Fine-tuning a small open-weight model to match it on that one narrow task can cut inference cost by an order of magnitude. This is distillation, and it is the strongest commercial argument for fine-tuning today.<\/li>\n\n\n\n<li><strong>Stable, distinctive behaviour.<\/strong> Tone, refusal patterns, house style, or a domain-specific response shape that must hold every single time.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">And three anti-signals, each a reliable way to waste a month:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Fine-tuning to add knowledge.<\/strong> It does not work well, it goes stale, and it removes your ability to cite sources.<\/li>\n\n\n\n<li><strong>Fine-tuning to stop hallucinations.<\/strong> A fine-tuned model tends to hallucinate more confidently, which is worse than before.<\/li>\n\n\n\n<li><strong>Fine-tuning without evaluation.<\/strong> If you cannot measure it, you cannot tell a good checkpoint from a bad one, and you will ship on vibes.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>What Are Parameters in an LLM?<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before the techniques make sense, this needs to be concrete, because every efficiency method below is a statement about which parameters you touch.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An LLM is, mechanically, an enormous set of numbers arranged in matrices. Those numbers are the <strong>parameters<\/strong>, also called weights. During pretraining, the model reads vast amounts of text and repeatedly nudges each number so that its predictions of the next token get better. When training finishes, those settled numbers <em>are<\/em> the model. Everything it appears to know lives in them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Model names announce the count. A 7B model has about seven billion parameters, a 70B model about seventy billion. More parameters generally means more capacity and also more memory, more cost, and more latency.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two practical consequences matter here:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Memory.<\/strong> Parameters must be held in GPU memory to run or train. At 16-bit precision a parameter takes two bytes, so a 7B model needs roughly 14 GB just to hold the weights for inference. Training needs several times that, because optimiser states and gradients are stored alongside them. This is the single reason full fine-tuning is expensive, and the single reason LoRA exists.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Which parameters change.<\/strong> Full fine-tuning updates all of them. PEFT methods freeze nearly all of them and train a small addition. That distinction is the whole efficiency story.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Supervised Fine-Tuning (SFT): Teaching the Task<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">SFT is the foundation, and everything else is either an efficiency trick applied to it or a stage that follows it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You assemble a dataset of input and output pairs showing exactly what you want: a support ticket and its correct structured classification, a document and its correct extraction, a question and an answer in your house style. You continue training the model on these pairs. It learns to produce that shape of output.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is the entire concept. The difficulty is never the training loop, which is a handful of lines with <a href=\"https:\/\/huggingface.co\/docs\/trl\/index\">Hugging Face TRL<\/a> or a managed service like <a href=\"https:\/\/platform.openai.com\/docs\/guides\/supervised-fine-tuning\">OpenAI&#8217;s fine-tuning API<\/a>. The difficulty is the data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What separates a good SFT dataset from a bad one:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Consistency beats volume.<\/strong> Five hundred examples that follow the same convention beat five thousand that disagree with each other. Contradictory labels teach the model to be inconsistent, and it will learn that faithfully.<\/li>\n\n\n\n<li><strong>The outputs must be exemplary, not merely acceptable.<\/strong> The model imitates what you show it, including sloppiness.<\/li>\n\n\n\n<li><strong>Cover the edge cases you actually care about.<\/strong> A dataset of only clean, easy examples produces a model that fails exactly where you needed help.<\/li>\n\n\n\n<li><strong>Hold out a test set before you start.<\/strong> Not after. This is the step teams skip and then regret.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The failure mode to know by name is <strong>catastrophic forgetting<\/strong>: train hard enough on a narrow task and the model degrades at everything else. A model fine-tuned aggressively on JSON extraction can lose general conversational ability. PEFT methods substantially reduce this risk, because the original weights are never overwritten.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>PEFT and LoRA: Fine-Tuning You Can Actually Afford<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Parameter-efficient fine-tuning<\/strong> covers all the methods that adapt a large model by training a small number of parameters while the base stays frozen. <a href=\"https:\/\/huggingface.co\/docs\/peft\/index\">Hugging Face PEFT<\/a> is the standard implementation library. In practice, PEFT mostly means LoRA.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>How LoRA works<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/arxiv.org\/abs\/2106.09685\">LoRA, Low-Rank Adaptation<\/a>, rests on an observation: the <em>change<\/em> you need to make to a pretrained weight matrix to adapt it to a new task has low intrinsic rank. It can be approximated by multiplying two much smaller matrices together.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So instead of updating a large weight matrix directly, LoRA freezes it and injects two small trainable matrices beside it. Their product is added to the original output. You train only those two, which is a tiny fraction of the total parameter count. The <strong>rank<\/strong>, usually written r, sets how small: higher rank means more capacity and more parameters.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The original paper reports reducing trainable parameters by up to 10,000 times and GPU memory requirements by three times, relative to full fine-tuning of GPT-3 175B with Adam, with model quality on par or better. The practical consequences for a working engineer are what matter:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Training fits on hardware you can rent by the hour.<\/li>\n\n\n\n<li>Adapters are small files, often a few megabytes, so you can version, ship and roll them back easily.<\/li>\n\n\n\n<li>You can keep several task-specific adapters and swap them over one shared base model in memory.<\/li>\n\n\n\n<li>Because the base weights are untouched, catastrophic forgetting is far less likely.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The <a href=\"https:\/\/huggingface.co\/docs\/peft\/developer_guides\/lora\">PEFT LoRA guide<\/a> covers configuration, including which modules to target and how to pick a rank.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>QLoRA<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/arxiv.org\/abs\/2305.14314\">QLoRA<\/a> adds one idea: quantise the frozen base model to 4-bit precision, then train LoRA adapters on top of it. Since the base is frozen anyway, the precision loss is far less damaging than it would be during full training. The paper demonstrates fine-tuning a 65B model on a single 48 GB GPU while preserving full 16-bit fine-tuning task performance.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The rule of thumb: <strong>use LoRA when memory allows, use QLoRA when it does not.<\/strong> Starting at 16-bit removes quantisation as a variable when you are debugging why a run underperformed.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Method<\/strong><\/td><td><strong>What trains<\/strong><\/td><td><strong>Memory<\/strong><\/td><td><strong>Best for<\/strong><\/td><\/tr><tr><td>Full fine-tuning<\/td><td>Every parameter<\/td><td>Very high<\/td><td>Labs, base model creation, deep domain shift<\/td><\/tr><tr><td>LoRA<\/td><td>Small adapter matrices<\/td><td>Moderate<\/td><td>The default for most production work<\/td><\/tr><tr><td>QLoRA<\/td><td>Adapters over a 4-bit base<\/td><td>Low<\/td><td>Large models on constrained or single-GPU setups<\/td><\/tr><tr><td>Prompt tuning<\/td><td>Soft prompt vectors only<\/td><td>Very low<\/td><td>Lightweight task switching, narrow gains<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>RLHF, DPO and Alignment: Teaching Preference<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">SFT teaches the model what a correct answer looks like. Alignment teaches it which of several correct answers is <em>better<\/em>: more helpful, safer, shorter, better reasoned. You cannot express that with labelled pairs, because preference is comparative.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>RLHF<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Reinforcement learning from human feedback<\/strong> is the method that made instruction-following chat models work, introduced at scale in <a href=\"https:\/\/arxiv.org\/abs\/2203.02155\">OpenAI&#8217;s InstructGPT paper<\/a>. Three stages:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>SFT.<\/strong> Fine-tune on demonstration data to get a competent starting model.<\/li>\n\n\n\n<li><strong>Reward model.<\/strong> Humans rank multiple model outputs for the same prompt. A separate model is trained to predict those human preferences and output a score.<\/li>\n\n\n\n<li><strong>RL optimisation.<\/strong> The language model is optimised, typically with PPO, to maximise the reward model&#8217;s score, with a penalty for drifting too far from the SFT model.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The headline result is worth remembering because it reframes what scale buys you: human labellers preferred the outputs of a 1.3B-parameter InstructGPT model over those of the 175B GPT-3, despite roughly 100 times fewer parameters. Alignment, not size, was the difference.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">RLHF is also genuinely difficult. Three models in play, an unstable RL loop, expensive human annotation, and reward hacking where the model learns to game the scorer rather than improve. For most engineering teams it is not the right tool.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>DPO, the practical alternative<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/arxiv.org\/abs\/2305.18290\">Direct Preference Optimization<\/a> delivers the same goal with far less machinery. Its insight is that the language model can be treated as its own implicit reward model, so you can optimise directly on preference pairs with a simple classification-style loss. No separate reward model, no RL loop.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You still need preference data, meaning pairs of a preferred and rejected response, but the training is stable and computationally light. <strong>For nearly every team, DPO is where preference alignment should start.<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>GRPO and verifiable rewards<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For tasks where correctness can be checked automatically, such as maths, code that compiles and passes tests, or schema-valid output, <strong>Group Relative Policy Optimization<\/strong>, introduced in the <a href=\"https:\/\/arxiv.org\/abs\/2402.03300\">DeepSeekMath paper<\/a>, is the current direction of travel.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It drops the separate value model and scores a group of sampled answers relative to each other, using a programmatic grader instead of human labels. If your task has a right answer a script can verify, this avoids the annotation bottleneck entirely.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Method<\/strong><\/td><td><strong>Data required<\/strong><\/td><td><strong>Complexity<\/strong><\/td><td><strong>When to use<\/strong><\/td><\/tr><tr><td>SFT<\/td><td>Input and output pairs<\/td><td>Low<\/td><td>Always first<\/td><\/tr><tr><td>DPO<\/td><td>Preference pairs<\/td><td>Low to moderate<\/td><td>Default for alignment<\/td><\/tr><tr><td>RLHF (PPO)<\/td><td>Rankings plus reward model<\/td><td>High<\/td><td>Research teams, frontier-scale work<\/td><\/tr><tr><td>GRPO<\/td><td>Automated verifier<\/td><td>Moderate<\/td><td>Maths, code, structured output<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>What Fine-Tuning Costs<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The frameworks are free. Your costs are GPU hours, data preparation, and maintenance, and engineers consistently underestimate the last two.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a typical LoRA or QLoRA run on a 7B to 13B model with a few thousand examples, the compute is usually measured in single-digit GPU hours on a rented A100 or equivalent. At Indian cloud and GPU-rental rates, a single experiment commonly lands in the low thousands of rupees. That is the number that surprises people: <strong>the compute is rarely the expensive part.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The real costs are elsewhere:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Data curation.<\/strong> Building and cleaning a few thousand high-quality examples is days of skilled work, and it is the majority of the effort on almost every project.<\/li>\n\n\n\n<li><strong>Evaluation.<\/strong> Building the harness that tells you whether the checkpoint is better.<\/li>\n\n\n\n<li><strong>Iteration.<\/strong> You will not get it right on run one. Budget for five to ten runs.<\/li>\n\n\n\n<li><strong>Maintenance.<\/strong> This is the cost nobody plans for. A fine-tuned model is a product with a lifecycle. Base models get deprecated, your data distribution drifts, and someone owns retraining forever. Before you start, decide who that is.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Verify current per-hour GPU pricing with your provider before budgeting; rates move quickly and vary widely between regions and providers.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>A Realistic First Project<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If you want to actually learn this rather than read about it, do a small version end to end:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Pick a narrow, checkable task: classifying support tickets into a fixed taxonomy, or extracting structured fields from documents.<\/li>\n\n\n\n<li>Write the evaluation first. A hundred held-out examples and a script that scores accuracy and schema validity.<\/li>\n\n\n\n<li>Get a prompt-only baseline. Record the number. This is what you must beat.<\/li>\n\n\n\n<li>Build 500 to 1,000 clean training examples.<\/li>\n\n\n\n<li>Run QLoRA on a small open-weight model using PEFT and TRL. <a href=\"https:\/\/docs.unsloth.ai\/\">Unsloth<\/a> is a good option for faster, lower-memory runs.<\/li>\n\n\n\n<li>Score the checkpoint against the same evaluation. Compare honestly against the baseline.<\/li>\n\n\n\n<li>Write down the cost per thousand inferences for both, then decide whether the fine-tune actually earned its keep.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Step 7 is the one that makes you an engineer rather than a tutorial-follower. Roughly half the time the honest answer is that the prompt was good enough, and reaching that conclusion with evidence is a genuinely valuable outcome. More hands-on builds of this kind are in\u00a0 gen ai projects list.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Where This Sits in an AI Engineering Career<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Fine-tuning is a late-stage skill, and treating it as an entry point is why many self-taught roadmaps stall. Interviewers do not usually ask whether you can run a LoRA script; that is a documented procedure. They ask whether you knew fine-tuning was the right answer, what you tried first, and how you proved it worked. The judgment is the skill.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The way to build that judgment is to run the decision honestly once, end to end, and be willing to conclude that the prompt won. Engineers who have done that can tell you what fine-tuning costs, when it pays, and how they proved it. Engineers who have only read about it reach for it first and find out later.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Conclusion<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The useful summary of fine tuning LLM work in 2026 is narrower than the hype suggests. Fine-tuning is for form, not facts. Climb the ladder in order: prompt, retrieve, add tools, and only then train. When you do train, use LoRA or QLoRA rather than full fine-tuning, run SFT before any alignment step, and reach for DPO rather than RLHF unless you have a research team.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Above all, build the evaluation harness before the training script. The engineers who get value from fine-tuning are not the ones who know the most about optimisers. They are the ones who can prove, with numbers, that the fine-tuned model beats the prompt it replaced. Everyone else is just spending GPU hours to feel productive.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>FAQ<\/strong><\/h2>\n\n\n\n<div class=\"schema-faq wp-block-yoast-faq-block\"><div class=\"schema-faq-section\" id=\"faq-question-1787667103132\"><strong class=\"schema-faq-question\"><strong>Is fine tuning an LLM better than RAG?<\/strong><\/strong> <p class=\"schema-faq-answer\">Neither is better; they solve different problems. RAG changes what the model can see at query time and is correct for knowledge that changes or must be cited. Fine-tuning changes how the model behaves every time and is correct for format, tone and task specialisation. Most mature production systems use both: retrieval for facts, a light fine-tune for behaviour.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667140214\"><strong class=\"schema-faq-question\"><strong>What is LoRA in simple terms?<\/strong><\/strong> <p class=\"schema-faq-answer\">LoRA freezes the model&#8217;s existing weights and adds two small trainable matrices alongside them, training only those. Because the added matrices are tiny relative to the model, you get most of the benefit of fine-tuning at a fraction of the memory and cost, and you can swap adapters in and out over one shared base model.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667159630\"><strong class=\"schema-faq-question\"><strong>Do I need RLHF for my application?<\/strong><\/strong> <p class=\"schema-faq-answer\">Almost certainly not. RLHF requires a reward model, an RL training loop and substantial human annotation. If you need preference alignment, start with DPO, which reaches a similar goal with a simple training loop and no reward model. If your task has an automatically verifiable right answer, look at GRPO instead.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667183517\"><strong class=\"schema-faq-question\"><strong>How many examples do I need to fine-tune an LLM?<\/strong> <\/strong> <p class=\"schema-faq-answer\">For a narrow task with LoRA, a few hundred to a few thousand high-quality examples is a workable range, and consistency matters more than count. Below a few hundred, the model mostly learns your labelling noise. The old rule that you need tens of thousands of examples came from full fine-tuning and no longer applies.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667193292\"><strong class=\"schema-faq-question\"><strong>Can fine-tuning teach a model new facts?<\/strong><\/strong> <p class=\"schema-faq-answer\">Technically yes, reliably no. Facts learned this way are hard to update, impossible to cite, and easy to overwrite in the next training run. If the model needs to know something, retrieve it. If the model needs to <em>behave<\/em> a certain way about something, fine-tune it.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667212183\"><strong class=\"schema-faq-question\"><strong>What is the difference between PEFT and LoRA?<\/strong><\/strong> <p class=\"schema-faq-answer\">PEFT is the category, LoRA is the most popular method within it. Other PEFT methods include prompt tuning and prefix tuning. In practice, when a team says PEFT they usually mean LoRA or QLoRA.<\/p> <\/div> <div class=\"schema-faq-section\" id=\"faq-question-1787667224403\"><strong class=\"schema-faq-question\"><strong>Will fine-tuning make my model worse at other tasks?<\/strong><\/strong> <p class=\"schema-faq-answer\">It can. Full fine-tuning on a narrow task risks catastrophic forgetting, where general capability degrades. LoRA and QLoRA reduce this substantially because the original weights stay frozen, but you should still evaluate general capability, not just your target task, before shipping<\/p> <\/div> <\/div>\n\n\n\n<h2 class=\"wp-block-heading\"><\/h2>\n","protected":false},"excerpt":{"rendered":"<p>Fine-tuning an LLM means updating a pretrained model&#8217;s weights on your own examples so the new behaviour is built into the model instead of described in a prompt. In 2026 the practical stack is three layers: supervised fine-tuning (SFT) teaches the model a task and an output format from labelled examples, parameter-efficient fine-tuning (PEFT) methods [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":208,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[20],"tags":[36],"class_list":["post-145","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","tag-machine-learning"],"blocksy_meta":{"styles_descriptor":{"styles":{"desktop":"","tablet":"","mobile":""},"google_fonts":[],"version":8}},"acf":{"reviewed_by":null},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.6 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog<\/title>\n<meta name=\"description\" content=\"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog\" \/>\n<meta property=\"og:description\" content=\"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/\" \/>\n<meta property=\"og:site_name\" content=\"Varsity Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-27T11:01:15+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-27T11:01:16+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png\" \/>\n\t<meta property=\"og:image:width\" content=\"741\" \/>\n\t<meta property=\"og:image:height\" content=\"487\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Varsity on Behalf of CEP IIT Delhi\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Varsity on Behalf of CEP IIT Delhi\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/\"},\"author\":{\"name\":\"Varsity on Behalf of CEP IIT Delhi\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#\\\/schema\\\/person\\\/7b5db9a94eddee529cd35968692d9c29\"},\"headline\":\"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers\",\"datePublished\":\"2026-08-27T11:01:15+00:00\",\"dateModified\":\"2026-08-27T11:01:16+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/\"},\"wordCount\":2970,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Screenshot-2026-08-26-145947.png\",\"keywords\":[\"Machine Learning\"],\"articleSection\":[\"Artificial Intelligence\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#respond\"]}]},{\"@type\":[\"WebPage\",\"FAQPage\"],\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/\",\"name\":\"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Screenshot-2026-08-26-145947.png\",\"datePublished\":\"2026-08-27T11:01:15+00:00\",\"dateModified\":\"2026-08-27T11:01:16+00:00\",\"description\":\"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#breadcrumb\"},\"mainEntity\":[{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667103132\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667140214\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667159630\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667183517\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667193292\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667212183\"},{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667224403\"}],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#primaryimage\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Screenshot-2026-08-26-145947.png\",\"contentUrl\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Screenshot-2026-08-26-145947.png\",\"width\":741,\"height\":487,\"caption\":\"Fine-Tuning- LLMs- SFT- LoRA- and -RLHF.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/\",\"name\":\"Varsity Blog\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#organization\",\"name\":\"Varsity Blog\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/varsity-logo.png\",\"contentUrl\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/varsity-logo.png\",\"width\":275,\"height\":64,\"caption\":\"Varsity Blog\"},\"image\":{\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/company\\\/varsity-by-interviewbit\\\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/#\\\/schema\\\/person\\\/7b5db9a94eddee529cd35968692d9c29\",\"name\":\"Varsity on Behalf of CEP IIT Delhi\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g\",\"caption\":\"Varsity on Behalf of CEP IIT Delhi\"},\"description\":\"Varsity by InterviewBit, in collaboration with CEP IIT Delhi, creates industry-relevant learning programmes designed to help learners build practical, in-demand skills. Through this author profile, we publish articles that complement our courses covering curriculum-aligned topics, foundational concepts, emerging trends, and advanced insights. Our goal is to help learners deepen their understanding beyond the classroom and apply their knowledge confidently in real-world contexts.\",\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/author\\\/varsity-on-behalf-of-cep-iit-delhi\\\/\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667103132\",\"position\":1,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667103132\",\"name\":\"Is fine tuning an LLM better than RAG?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Neither is better; they solve different problems. RAG changes what the model can see at query time and is correct for knowledge that changes or must be cited. Fine-tuning changes how the model behaves every time and is correct for format, tone and task specialisation. Most mature production systems use both: retrieval for facts, a light fine-tune for behaviour.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667140214\",\"position\":2,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667140214\",\"name\":\"What is LoRA in simple terms?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"LoRA freezes the model's existing weights and adds two small trainable matrices alongside them, training only those. Because the added matrices are tiny relative to the model, you get most of the benefit of fine-tuning at a fraction of the memory and cost, and you can swap adapters in and out over one shared base model.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667159630\",\"position\":3,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667159630\",\"name\":\"Do I need RLHF for my application?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Almost certainly not. RLHF requires a reward model, an RL training loop and substantial human annotation. If you need preference alignment, start with DPO, which reaches a similar goal with a simple training loop and no reward model. If your task has an automatically verifiable right answer, look at GRPO instead.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667183517\",\"position\":4,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667183517\",\"name\":\"How many examples do I need to fine-tune an LLM?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"For a narrow task with LoRA, a few hundred to a few thousand high-quality examples is a workable range, and consistency matters more than count. Below a few hundred, the model mostly learns your labelling noise. The old rule that you need tens of thousands of examples came from full fine-tuning and no longer applies.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667193292\",\"position\":5,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667193292\",\"name\":\"Can fine-tuning teach a model new facts?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Technically yes, reliably no. Facts learned this way are hard to update, impossible to cite, and easy to overwrite in the next training run. If the model needs to know something, retrieve it. If the model needs to <em>behave<\\\/em> a certain way about something, fine-tune it.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667212183\",\"position\":6,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667212183\",\"name\":\"What is the difference between PEFT and LoRA?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"PEFT is the category, LoRA is the most popular method within it. Other PEFT methods include prompt tuning and prefix tuning. In practice, when a team says PEFT they usually mean LoRA or QLoRA.\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"},{\"@type\":\"Question\",\"@id\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667224403\",\"position\":7,\"url\":\"https:\\\/\\\/www.interviewbit.com\\\/varsity\\\/blog\\\/fine-tuning-llms\\\/#faq-question-1787667224403\",\"name\":\"Will fine-tuning make my model worse at other tasks?\",\"answerCount\":1,\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"It can. Full fine-tuning on a narrow task risks catastrophic forgetting, where general capability degrades. LoRA and QLoRA reduce this substantially because the original weights stay frozen, but you should still evaluate general capability, not just your target task, before shipping\",\"inLanguage\":\"en-US\"},\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog","description":"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/","og_locale":"en_US","og_type":"article","og_title":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog","og_description":"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.","og_url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/","og_site_name":"Varsity Blog","article_published_time":"2026-08-27T11:01:15+00:00","article_modified_time":"2026-08-27T11:01:16+00:00","og_image":[{"width":741,"height":487,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png","type":"image\/png"}],"author":"Varsity on Behalf of CEP IIT Delhi","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Varsity on Behalf of CEP IIT Delhi","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#article","isPartOf":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/"},"author":{"name":"Varsity on Behalf of CEP IIT Delhi","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#\/schema\/person\/7b5db9a94eddee529cd35968692d9c29"},"headline":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers","datePublished":"2026-08-27T11:01:15+00:00","dateModified":"2026-08-27T11:01:16+00:00","mainEntityOfPage":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/"},"wordCount":2970,"commentCount":0,"publisher":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#organization"},"image":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#primaryimage"},"thumbnailUrl":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png","keywords":["Machine Learning"],"articleSection":["Artificial Intelligence"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#respond"]}]},{"@type":["WebPage","FAQPage"],"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/","name":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers - Varsity Blog","isPartOf":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#primaryimage"},"image":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#primaryimage"},"thumbnailUrl":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png","datePublished":"2026-08-27T11:01:15+00:00","dateModified":"2026-08-27T11:01:16+00:00","description":"Learn how to fine-tune LLMs with SFT, LoRA, and RLHF. Understand key techniques, workflows, trade-offs, and practical considerations for working engineers.","breadcrumb":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#breadcrumb"},"mainEntity":[{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667103132"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667140214"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667159630"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667183517"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667193292"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667212183"},{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667224403"}],"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#primaryimage","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png","contentUrl":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-26-145947.png","width":741,"height":487,"caption":"Fine-Tuning- LLMs- SFT- LoRA- and -RLHF.png"},{"@type":"BreadcrumbList","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.interviewbit.com\/varsity\/blog\/"},{"@type":"ListItem","position":2,"name":"Fine-Tuning LLMs: SFT, LoRA and RLHF Explained for Working Engineers"}]},{"@type":"WebSite","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#website","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/","name":"Varsity Blog","description":"","publisher":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.interviewbit.com\/varsity\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#organization","name":"Varsity Blog","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/varsity-logo.png","contentUrl":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-content\/uploads\/2026\/08\/varsity-logo.png","width":275,"height":64,"caption":"Varsity Blog"},"image":{"@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.linkedin.com\/company\/varsity-by-interviewbit\/"]},{"@type":"Person","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/#\/schema\/person\/7b5db9a94eddee529cd35968692d9c29","name":"Varsity on Behalf of CEP IIT Delhi","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/bdb12370d043ff980b2b8f1ccb67f5a1f0e333aaca46cc35358b1af8b1d98334?s=96&d=mm&r=g","caption":"Varsity on Behalf of CEP IIT Delhi"},"description":"Varsity by InterviewBit, in collaboration with CEP IIT Delhi, creates industry-relevant learning programmes designed to help learners build practical, in-demand skills. Through this author profile, we publish articles that complement our courses covering curriculum-aligned topics, foundational concepts, emerging trends, and advanced insights. Our goal is to help learners deepen their understanding beyond the classroom and apply their knowledge confidently in real-world contexts.","url":"https:\/\/www.interviewbit.com\/varsity\/blog\/author\/varsity-on-behalf-of-cep-iit-delhi\/"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667103132","position":1,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667103132","name":"Is fine tuning an LLM better than RAG?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"Neither is better; they solve different problems. RAG changes what the model can see at query time and is correct for knowledge that changes or must be cited. Fine-tuning changes how the model behaves every time and is correct for format, tone and task specialisation. Most mature production systems use both: retrieval for facts, a light fine-tune for behaviour.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667140214","position":2,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667140214","name":"What is LoRA in simple terms?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"LoRA freezes the model's existing weights and adds two small trainable matrices alongside them, training only those. Because the added matrices are tiny relative to the model, you get most of the benefit of fine-tuning at a fraction of the memory and cost, and you can swap adapters in and out over one shared base model.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667159630","position":3,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667159630","name":"Do I need RLHF for my application?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"Almost certainly not. RLHF requires a reward model, an RL training loop and substantial human annotation. If you need preference alignment, start with DPO, which reaches a similar goal with a simple training loop and no reward model. If your task has an automatically verifiable right answer, look at GRPO instead.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667183517","position":4,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667183517","name":"How many examples do I need to fine-tune an LLM?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"For a narrow task with LoRA, a few hundred to a few thousand high-quality examples is a workable range, and consistency matters more than count. Below a few hundred, the model mostly learns your labelling noise. The old rule that you need tens of thousands of examples came from full fine-tuning and no longer applies.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667193292","position":5,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667193292","name":"Can fine-tuning teach a model new facts?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"Technically yes, reliably no. Facts learned this way are hard to update, impossible to cite, and easy to overwrite in the next training run. If the model needs to know something, retrieve it. If the model needs to <em>behave<\/em> a certain way about something, fine-tune it.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667212183","position":6,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667212183","name":"What is the difference between PEFT and LoRA?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"PEFT is the category, LoRA is the most popular method within it. Other PEFT methods include prompt tuning and prefix tuning. In practice, when a team says PEFT they usually mean LoRA or QLoRA.","inLanguage":"en-US"},"inLanguage":"en-US"},{"@type":"Question","@id":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667224403","position":7,"url":"https:\/\/www.interviewbit.com\/varsity\/blog\/fine-tuning-llms\/#faq-question-1787667224403","name":"Will fine-tuning make my model worse at other tasks?","answerCount":1,"acceptedAnswer":{"@type":"Answer","text":"It can. Full fine-tuning on a narrow task risks catastrophic forgetting, where general capability degrades. LoRA and QLoRA reduce this substantially because the original weights stay frozen, but you should still evaluate general capability, not just your target task, before shipping","inLanguage":"en-US"},"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/posts\/145","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/comments?post=145"}],"version-history":[{"count":2,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/posts\/145\/revisions"}],"predecessor-version":[{"id":237,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/posts\/145\/revisions\/237"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/media\/208"}],"wp:attachment":[{"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/media?parent=145"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/categories?post=145"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.interviewbit.com\/varsity\/blog\/wp-json\/wp\/v2\/tags?post=145"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}