Skip to content
IronbridgeAI
All insights

Custom AI Apps

RAG vs Agentic AI: Choosing Your AI App Architecture

By Jordan SolenderUpdated Sep 30, 20268 min read

In RAG vs agentic AI, the choice depends on the job. Use retrieval augmented generation (RAG) when the app’s job is to answer questions from your documents and data. Use an agent when the job is to do something: update a record, send an email, schedule a meeting. Most production apps combine them, with an agent that uses retrieval as one of its tools.

This guide covers how each pattern works, how to choose, how to pick an LLM, and how to “train” a chatbot on your SOPs and documents, which is really a retrieval project. For the wider build process, see how to build a custom AI web app.

What RAG and agents actually are

RAG and agents solve different problems. RAG gives a model the right context. An agent gives a model the ability to act.

RAG works in two steps. When a question comes in, the app searches your content (documents, tickets, records) for the most relevant passages. It then passes those passages to the model with the question, and the model answers from them, ideally with citations. The model’s knowledge stays the same; you are handing it the right pages at the right moment.

An agent is a model in a loop with tools. It reads a goal, decides which tool to call (search the CRM, draft an email, create a task), looks at the result, and decides the next step until the goal is met or it hands off to a person. The tools are functions your code defines, so your code still controls what the agent is allowed to do.

RAG vs agents compared

The simplest test: if the output is text a person reads, start with RAG. If the output is a change in another system, you need an agent.

RAG Agent
Job Answer, summarize, find Complete multi-step tasks
Output Text with citations Actions in other systems, plus text
Typical uses Policy Q&A, contract lookup, support knowledge base Update CRM, route email, book meetings, draft follow-ups
Main risk Wrong or stale source documents Wrong action taken with real consequences
Main control Document quality and permission filters Tool permissions, approvals, audit logs
Testing Question and answer eval set Task scenarios with expected actions
Build effort Lower Higher

How RAG and agents combine

Most useful business apps are agents that use RAG as one of their tools. The agent decides what to do. Retrieval gives it grounded context so it does it well.

A real example: an IT solutions provider replaced its Salesforce setup with an AI-native revenue portal. Retrieval-style work mined 1,630 recorded sales-call transcripts for renewal dates, pain points and incumbent vendors. Agent-style work routes every email and meeting to the right deal automatically, and a nightly AI analyst drafts account plays that a person approves. Neither pattern alone would have covered it.

A second example: a coaching and media company’s client portal generates AI call recaps and an AI-written weekly brief. That is mostly retrieval and summarization over the company’s own recordings and records, with a light scheduled workflow around it. No autonomous actions were needed, so none were built.

How to choose your architecture

Pick the simplest pattern that does the job, then add autonomy only where it pays. Every step of autonomy adds testing and risk.

  1. Write the job as a sentence. “Answer staff questions about our return policy” is RAG. “When a lead replies, log it and draft a follow-up” is an agent.
  2. List the systems it must read from and write to. Read-only access points to RAG. Any write points to an agent with tools.
  3. Decide where a human approves. Anything that sends, deletes, pays or commits the company should start with a person clicking approve.
  4. Start with a fixed workflow before a free-roaming agent. Many “agents” work best as a defined sequence of steps with a model at a few decision points. Give the model open-ended choice only where the path truly varies.
  5. Write the eval set for that pattern. Questions and expected answers for RAG; scenarios and expected actions for an agent.

For agents that act inside your business systems, our AI agents service covers the tool design and guardrails in more depth.

How to choose an LLM for your app

Choose the model with an eval set, not a leaderboard. Model choice is one of the biggest cost levers in an AI app, but it matters less than teams expect once prompts, retrieval and evals are solid.

Build the eval set first

Before comparing models, write 20 to 50 real examples of the task with the output you would accept. Include the hard and messy cases, not only the clean ones. Run each candidate model against the set and score it. This turns “which model is best” into a measurable question for your job.

Start with a smaller model, upgrade where it fails

Every major provider offers tiers: Anthropic’s Claude Haiku, Sonnet and Opus, OpenAI’s smaller and full GPT models, and Google’s Gemini Flash and Pro. Smaller tiers are faster and much cheaper per token, and they handle many routine tasks such as classification, extraction and short drafts well. Try the smaller model first and move up only for the tasks your eval set shows it cannot handle.

Plan for more than one model

Most production apps route different tasks to different models: simple routing or tagging to a small model, complex reasoning or long documents to a larger one. Set up that routing early, even if only one model is used at first. It also gives you a fallback when one provider has an outage.

Other factors that matter

  • Data terms. Check retention and training policies on the API tier you use, especially for client or employee data.
  • Context length. Long contracts or transcripts may need a model with a large context window, or better retrieval so you send less.
  • Tool use quality. For agents, test how reliably each model calls tools with correct arguments.
  • Latency. Chat interfaces need fast first tokens; overnight batch jobs do not.

How to train an AI chatbot on your SOPs and documents

“Training” a chatbot on company documents almost never means retraining a model. It means giving a model reliable, permissioned access to your documents at the moment a question is asked. That is RAG, and it is cheaper, faster and easier to keep current than fine-tuning.

Why retrieval beats fine-tuning for SOPs

Fine-tuning bakes knowledge into model weights. When an SOP changes next month, you retrain. Retrieval reads the current document at question time, so updating a policy in your document system updates the assistant. Retrieval also lets the assistant cite its source, which is what makes a team trust it instead of quietly abandoning it.

Step 1: prepare the documents

This is where most of the work lives. Consolidate scattered docs, delete outdated versions, and give each document a clear title, owner and last-reviewed date. Contradictory SOPs produce contradictory answers, and no model fixes that.

Step 2: chunk and index sensibly

Split documents by section and heading, not by arbitrary character counts, so each retrieved passage makes sense on its own. Keep the document title and section name attached to each chunk. Store them in a vector store such as pgvector or Pinecone, often alongside keyword search, which catches exact terms like product codes.

Step 3: respect permissions

The assistant must inherit your existing access controls. HR files, compensation data and client contracts should surface only for people already allowed to read them. Filter at retrieval time based on who is asking, never after the model has already seen the content.

Step 4: evaluate before rollout

Write about fifty real questions your team actually asks, with correct answers and the source document for each. Score accuracy and citation quality. Agree on the pass bar with the document owners before you launch, and keep this set as your regression test for every document or prompt change.

Step 5: roll out with a feedback loop

Add thumbs up and down to every answer and review the negatives weekly. Nearly all early failures trace back to a missing, stale or badly written document rather than the model. Fixing the document fixes the answer.

Frequently asked questions

What is the difference between RAG and agentic AI?

RAG retrieves relevant content from your data and hands it to a model so answers are grounded and cited. Agentic AI gives a model tools and lets it take multi-step actions, such as updating records or sending messages. RAG answers; agents act.

Can an AI agent use RAG?

Yes, and most production agents do. Retrieval is one of the agent’s tools, used to look up policies, past conversations or account history before deciding what to do. The combination gives you grounded decisions and real actions.

Should I fine-tune a model or use RAG for company documents?

Use RAG for company knowledge that changes, such as SOPs, policies and product information. Fine-tuning is better suited to changing a model’s style or format on a narrow task, and it does not keep up with document edits. Most businesses never need fine-tuning.

Which LLM is best for a custom AI app?

There is no single best model. Build an eval set of real examples, test a small and a large model from two providers, and pick the cheapest one that passes. Route harder tasks to a stronger model only where the eval set shows you need it.

How do I stop an AI chatbot from making things up about our policies?

Ground every answer in retrieved documents, require citations, and tell the model to say it does not know when the documents do not cover the question. Then fix the source documents behind every wrong answer you find in feedback.

Is an agentic AI app riskier than a RAG app?

Yes, because agents change things in other systems. Limit each agent’s tools to what the job needs, check permissions in code, require approval for sensitive actions, and keep an audit log. See our custom AI applications page and custom AI web apps service for how we scope this.

If you are deciding between RAG, an agent or both for a specific workflow, book a strategy call and we will map the architecture with you.