RAG vs Fine-Tuning: A Developer Guide
I keep getting the same question from clients building anything with an LLM behind it: should this use RAG, or should we just fine-tune a model on our own data? I've shipped both approaches on real projects now — a support bot pulling from a changing knowledge base, and a classification tool fine-tuned on a client's own historical tickets — and the honest answer is that RAG vs fine-tuning isn't really a quality contest. It's a question about how often your data changes and how much you're willing to pay to update it.
What RAG actually buys you
Retrieval-Augmented Generation means your app pulls relevant chunks of your own data (docs, past tickets, product info) at request time and stuffs them into the prompt before calling the model. Nothing about the model itself changes. You update a vector store, and the next request already sees the new information — no retraining, no waiting, no separate ML pipeline.
// Minimal RAG lookup before calling the model
async function answerWithContext(question, pgPool, openai) {
const { rows: [{ embedding }] } = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: question,
}).then(r => ({ rows: [{ embedding: r.data[0].embedding }] }));
const { rows } = await pgPool.query(
`SELECT content FROM docs
ORDER BY embedding <-> $1
LIMIT 5`,
[embedding]
);
const context = rows.map(r => r.content).join('\n---\n');
return openai.chat.completions.create({
model: 'gpt-4o',
messages: [
{ role: 'system', content: `Answer only using this context:\n${context}` },
{ role: 'user', content: question },
],
});
}
This is roughly what I run for a client whose product docs change weekly. The whole thing updates by re-embedding changed rows — there's no model to retrain, which matters a lot when the underlying content is a moving target.
The tradeoff nobody mentions upfront: RAG's retrieval step adds a real cost per request, not just latency. Embedding the question, running the vector search, and padding the prompt with retrieved chunks all add up, and on a high-traffic endpoint that difference shows up on the OpenAI bill by the end of the month — it's not free just because there's no training step.
What fine-tuning actually buys you
Fine-tuning means training the model itself on examples of the exact input/output pattern you want, so it stops needing instructions and context stuffed into every prompt — it just learned the behavior. I did this for a client who needed incoming support tickets auto-classified into about 40 internal categories, using a few thousand historically-labeled tickets. RAG would have meant stuffing category definitions and examples into every single request; fine-tuning meant the model just knows the categories cold, with a shorter prompt and a cheaper, faster call every time.
The comparison that actually decides it
| RAG | Fine-tuning | |
|---|---|---|
| Data changes often | Handles it well — just re-embed | Needs a retrain, which costs time and money |
| Data is mostly static | Works, but overkill | Ideal — pay the training cost once |
| Needs to cite sources / "show your work" | Natural fit — you control what's retrieved | Hard — the model can't point back to a source doc |
| Output is a narrow, repeatable pattern (classification, specific tone) | Possible but verbose prompts | Ideal — the pattern becomes the model's default behavior |
| Setup cost | A vector DB and an embedding pipeline | A labeled dataset and a training run |
| Per-request cost | Higher — bigger prompts from injected context | Lower — shorter prompts once the model knows the pattern |
| Latency | Slightly higher — retrieval step adds a hop | Slightly lower — no retrieval step |
The framing I actually use with clients: if someone could ask "where did that answer come from," you want RAG. If the task is "do this exact thing the same way every time," fine-tuning wins, and I'd actually push back on a client's instinct to reach for RAG by default just because it's the trendier term right now — a lot of narrow classification tasks are a worse fit for RAG than people assume, because you end up re-explaining the same categories in every prompt instead of letting the model just know them.
A hybrid is usually the real answer
Most production systems I've actually built end up doing both: fine-tune the model lightly on your team's tone and output format, then use RAG to inject the specific facts that change. The fine-tuning handles "sound like us and answer in this shape," and the RAG layer handles "here's today's actual product data." Treating it as an either/or question is where most teams waste time deciding — it's rarely a single clean choice in a real system.
Frequently Asked Questions
Can I fine-tune AND use RAG on the same model? Yes — fine-tune for tone/format/behavior, then still inject retrieved context at request time for facts. They solve different problems and stack fine.
Is fine-tuning worth it for a small dataset, like a few hundred examples? Usually not for OpenAI-style fine-tuning — you generally want at least a thousand or so clean, consistent examples before the model reliably picks up the pattern. Under that, prompt engineering or RAG alone usually gets you further for less effort.
Does RAG eliminate hallucinations completely? No. It reduces them significantly by grounding answers in real retrieved text, but the model can still misread or over-generalize from what it retrieved — always instruct it explicitly to say "I don't know" rather than guess when nothing relevant comes back.
How often should I re-embed my RAG data? Depends entirely on how often the source changes — I run it on a schedule tied to the actual update cadence of the source (hourly for a changing support queue, nightly for slower-moving product docs), not a fixed interval picked arbitrarily.
Related posts
Fix Invalid JSON from LLM API Responses
Why LLM APIs return invalid JSON and how to fix it for good: structured output modes, schema validation, and a safe parse-and-repair fallback in Node.js.
AI Prompt to Document an Undocumented API
Turn a day of writing API docs by hand into an hour of review with the right AI prompt structure for documenting undocumented REST endpoints.
Debug a Stack Trace Faster With AI
How to debug a Node.js stack trace with AI the right way: the exact prompt structure that gives real fixes instead of generic null-check guesses.
AI Refactoring: Breaking Up a God Class
A real AI refactoring workflow for breaking up a legacy god class safely: the prompts that work, and the ones that quietly break production.
0 Comments
No comments yet — be the first to share your thoughts.