AI · October 8, 2026 · 5 min read

RAG, fine-tuning or agents? How to choose the right approach

When companies start adding AI to their products or processes, three terms come up quickly: RAG, fine-tuning and agents. They're often presented as competing options. In practice they solve different problems, and many useful systems use more than one.

Here's what each one actually does, when it fits, and the order that usually makes sense to try them.

Start with the simplest option: a good prompt

Before any of the three, try a well-written prompt on a capable model. Give it clear instructions, a few examples of good output and the information it needs for the task. Modern models handle a wide range of jobs this way, from classifying requests to drafting replies and extracting fields from documents.

If a prompt with your real examples gets close to good enough, you may not need anything more complex. If it falls short, the way it falls short tells you which approach to reach for next.

RAG: give the model your information

Retrieval-augmented generation, or RAG, means searching your own content for the passages relevant to a question and including them in the prompt. The model then answers from that material instead of from its general knowledge.

  • Fits when the model needs facts it doesn't have: your docs, policies, product data, contracts or past tickets.
  • Fits when the information changes often, since you update the content, not the model.
  • Makes answers traceable, because you can show which sources were used.
  • Depends heavily on retrieval quality. If the search returns the wrong passages, the answer will be wrong too.

Most of the work in RAG is not the model. It's cleaning up the content, splitting it sensibly, getting search to return the right passages and keeping access rules intact, so people only see answers drawn from documents they're allowed to see. Our post on AI support triage walks through a RAG setup for help center content.

Fine-tuning: change how the model behaves

Fine-tuning means training an existing model further on your own examples, so it learns a pattern: a format, a tone, a classification scheme or a narrow task. You provide many input and output pairs, and the result is a customized version of the model.

  • Fits when you need consistent behavior that's hard to describe in a prompt, like a very specific output style.
  • Fits when a smaller fine-tuned model can replace a larger one for a narrow, high-volume task, reducing cost and response time.
  • Is a poor way to teach facts. Knowledge added through fine-tuning is hard to update and hard to verify. Use RAG for facts.
  • Needs a solid set of good examples, plus an evaluation set to prove the tuned model is actually better.

Fine-tuning also creates something to maintain. When a better base model comes out, you may need to tune again, and not every model or provider offers it on the same terms. For many teams, a strong prompt with examples gets most of the benefit with less upkeep.

Agents: let the model take steps

An agent is a setup where the model can decide to call tools, such as searching a database, reading a file, calling an API or sending a message, look at the results and decide what to do next, in a loop, until the task is done.

  • Fits when the task needs several steps that depend on each other, and the steps vary from case to case.
  • Fits when the work spans multiple systems, like checking an order, looking up the customer and drafting a response.
  • Adds risk with every tool it can use, especially tools that change data, spend money or contact people.
  • Is harder to test, slower and more expensive per task than a single model call, because it makes many calls.

Agents need guardrails that match their power: the narrowest set of tools, read-only access where possible, limits on steps and spending, logging of every action and a person approving anything that can't be undone. If the steps are always the same, a fixed workflow that calls the model at certain points is simpler and more reliable than an agent.

How they fit together

The three approaches aren't exclusive. A typical internal assistant might combine them.

  • An agent decides which steps to take for a request.
  • One of its tools is RAG search over company documents, filtered by the user's permissions.
  • A small fine-tuned model classifies incoming requests cheaply before the main model sees them.
  • Plain code handles everything that doesn't need judgment, like validation, formatting and saving results.

That last point matters. The more of a system that's ordinary, tested code, the easier it is to trust and maintain. Use the model for the parts that genuinely need language understanding.

What to try first

  1. Write down the task and collect real examples with expected outputs, so you can measure every option the same way.
  2. Try a strong prompt with a few examples. Measure it.
  3. If the model lacks your information, add RAG. Measure again.
  4. If the task needs several varying steps across systems, add tools, starting with read-only ones, and keep a person in the loop.
  5. If behavior is still inconsistent, or cost at high volume is a problem, consider fine-tuning a smaller model on examples you've collected along the way.

Each step adds cost, complexity and maintenance. Move to the next only when the measurements show the simpler approach isn't enough.

Deeraf helps companies choose and build the right approach for each AI integration, on infrastructure they control, starting with a Tech Check or going straight to a Build Sprint.

Keep reading

Want a second pair of eyes on your app?

Book a Tech Check