Self-Hosted Agentic RAG Guide 2026: Patterns, Tools & Implementation | AI Tutorial
Traditional RAG retrieves once, stuffs chunks into context, and hopes the LLM can answer. For simple FAQ-style questions, that works fine. But when someone asks your AI: “Compare retention rates between our enterprise and startup plans last quarter, factoring in the March compliance changes” — a single-shot retrieval doesn’t cut it. The answer requires multiple sources, computation, and multi-step reasoning that a passive retrieval pipeline simply cannot handle.
Agentic RAG fixes this. Instead of one retrieval call at the start, an LLM agent takes control of the retrieval process — planning what to look up, evaluating results, refining queries, and looping until it has enough evidence. The model becomes a research analyst, not a lookup tool.
This guide covers what agentic RAG actually is, the five core patterns that work in production, how to build it with self-hosted tools in 2026, and the tradeoffs you need to understand before committing.
What This Article Covers
- How agentic RAG differs from classic RAG and why the distinction matters
- Five canonical agentic retrieval patterns with real tradeoffs
- Self-hosted tool choices: Ollama, LangGraph, LlamaIndex, Qdrant
- Step-by-step implementation sketch
- Iteration budgets and cost control
- Common failure modes and how to avoid them
- Decision matrix: when agentic beats classic RAG
Classic RAG vs Agentic RAG: The Fundamental Shift
Classic RAG is a pipeline. Query goes in, top-k chunks come back, LLM answers. The structure is fixed — only the content changes. It works well for factual lookups where the answer lives in a single document section. It fails silently when the chunks don’t contain the answer even though the answer exists elsewhere in your corpus.
Agentic RAG flips this. Retrieval becomes a tool the agent can invoke whenever it decides more context would help. The agent reads what was retrieved, evaluates whether it answers the question, and chooses its next move: re-retrieve with a refined query, decompose into sub-queries, triangulate across corpora, or commit to an answer. The loop runs until the model signals confidence or a stop condition fires.
The key difference: Classic RAG treats retrieval as preprocessing. Agentic RAG treats retrieval as a tool call, under agent control.
Comparison: Classic RAG vs Agentic RAG
| Dimension | Classic RAG | Agentic RAG |
|---|---|---|
| Retrieval calls | 1 (fixed) | 2–7 (agent decides) |
| Latency | 1–3 seconds | 10–60 seconds |
| Token cost | Baseline | 3–10x baseline |
| Best for | FAQ, chat, scoped lookups | Research, synthesis, multi-hop |
| Failure mode | Misses context silently | Runaway cost, loop conditions |
| Query refinement | None | Agent reformulates queries |
| Self-evaluation | None | Agent judges retrieval quality |
| Reasoning | Single inference | Multi-step chain-of-thought |
Classic RAG fails when the query requires multiple data sources, comparative analysis, or computation. Agentic RAG handles those cases — at the cost of higher latency and token usage. The craft is knowing which workload shape you have and routing accordingly.
The Five Core Agentic RAG Patterns
These five patterns cover most production use cases. They range from simple (iterative retrieval) to complex (cross-corpus triangulation). Start with pattern 1 if you’re new to agentic RAG, and layer in others as your queries demand.
Pattern 1: Iterative Retrieval with Reflection
The foundational pattern. The agent retrieves, reads, critiques what it found, and decides whether to re-retrieve with a refined query.
The loop:
- Retrieve: query the vector store with the current query formulation
- Read: model consumes chunks and assesses relevance
- Critique: model articulates what it still doesn’t know, what contradicts, or where coverage is thin
- Refine: model rewrites the query to target the gap and loops back, or commits to an answer if confident
The critique step is the whole point — and where most implementations fail. A critique that just says “I need more information” is useless. The agent will re-retrieve with the same query and get the same chunks. A useful critique names what’s missing specifically: “The retrieved content covers 2024 pricing but not the 2026 update. I should retrieve for recent pricing announcements.” That specificity is what makes the next retrieval different from the last.
Practical tip: Force a structured critique schema — a JSON object with fields for “answered,” “missing,” “contradictions,” and “proposed next query.” This prevents the agent from endlessly rephrasing the original query without making progress.
{"answered": ["base pricing structure", "tier definitions"], "missing": ["April 2026 price changes", "regional variations"], "contradictions": [], "next_query": "claude opus 4.7 pricing changes April 2026", "confidence": 0.62, "should_continue": true}
Best for: Queries where initial retrieval is partially correct but incomplete, or where the user query uses different vocabulary than the corpus.
Pattern 2: Query Decomposition
Some questions cannot be answered by a single retrieval no matter how sharp the query is. “How does our latency compare to competitors in the EU after the March compliance changes?” is three retrievals stacked: your current latency, competitor latency, and the March compliance changes. Decomposition breaks the hard query into a tree of sub-queries, retrieves each, and synthesizes the result.
The process:
- Agent reads the original query and plans 2–6 sub-queries
- Each leaf sub-query runs through classic or iterative RAG
- Agent synthesizes leaf results into the final answer, flagging any sub-query that failed
- If a sub-query failed, agent decides whether to re-plan, retry, or answer with partial coverage
Good decomposition has two rules: Sub-queries must be independently retrievable — each one must make sense without context from the others. And decomposition trees should be flat where possible. A three-level tree burns 3x the retrievals of a two-level tree for marginal coverage gains. Cap at 6 leaves and one level of nesting unless the query is genuinely hierarchical.
Parallel vs Sequential: Independent sub-queries should run in parallel — three retrievals at 2 seconds each takes 2 seconds parallel and 6 seconds sequential. The agent decides which is which during planning.
Best for: Multi-hop questions, comparative analysis, questions spanning multiple documents or data sources.
Pattern 3: Hypothesis-Driven Retrieval
Inverted retrieval. Instead of “find me relevant content,” the agent forms a hypothesis about the answer, then retrieves specifically to confirm or deny it. This pattern is useful when the task is less “summarize the corpus” and more “is X true?”
The process:
- Agent articulates a specific, falsifiable claim: “I hypothesize X is true because Y”
- Agent retrieves content that would support or contradict the hypothesis
- If evidence is mixed, agent sharpens the query and retrieves again
- If evidence is conclusive, agent answers
Why hypotheses beat open queries: Vector search returns semantically similar content, which isn’t the same as answer-relevant content. A hypothesis gives the agent a concrete target to evaluate against, sharpening both the retrieval query and reading comprehension. In evaluations on legal and medical research tasks, hypothesis-driven retrieval converges in fewer iterations than open-ended iterative retrieval — because the agent stops once the hypothesis is settled rather than endlessly looking for “more context.”
The failure mode: Hypothesis-driven retrieval fails when the hypothesis is wrong in a way the agent cannot detect. It searches for evidence, finds it, confirms, and moves on — missing the better answer elsewhere. Mitigation: require the agent to also search for disconfirming evidence before committing, and flag queries where the confirming-to-disconfirming ratio is suspiciously high as needing triangulation.
Best for: Research tasks where the question is “is X true?”, investigation-style queries, legal and medical review.
Pattern 4: Cross-Corpus Triangulation
Run the same query against multiple retrieval sources and fuse the results. When independent corpora agree, confidence rises. When they disagree, the disagreement itself is useful signal — either the sources differ in scope, one is out of date, or the question is genuinely contested.
Common source combinations:
- Vector store over internal documents for semantic recall
- Knowledge graph for structured entity and relationship queries
- Web search for recency and public-facing claims
- SQL or analytics tools for exact numeric lookups
- Secondary vector index over a different corpus for domain-specific coverage
Fusion strategies: The simplest is presenting all retrieved content to the model with source tags and letting it reconcile. For larger sets, reciprocal rank fusion (RRF) and learned re-ranking models outperform naive concatenation. The agent picks chunks across sources up to a token budget, weighted by source reliability and rank.
Confidence from agreement: When three independent corpora return the same answer, confidence should be higher than when one returns the answer and two return nothing. The agent should surface that explicitly: “All three sources agree on X” versus “Only the internal wiki mentions X.”
Best for: Research synthesis, verification tasks, queries where source reliability varies.
Pattern 5: Evidence-Weighted Synthesis
When retrieval surfaces conflicting information, this pattern weighs each piece of evidence by source reliability, recency, and specificity, then produces a synthesis that reflects those weights.
Evidence weighting guide:
| Evidence Type | Typical Weight | Primary Signal |
|---|---|---|
| Primary source document | High | Author authority, publication date |
| Internal wiki / knowledge base | Medium-high | Last-updated timestamp, review status |
| Aggregated / summarized content | Medium | Underlying sources, synthesis date |
| Community / forum content | Low-medium | Engagement signals, corroboration |
| Outdated cached content | Low | Used only if nothing else available |
Citation integrity in synthesis: The hard part is keeping citations correct when the final answer blends multiple sources. Tag each piece of evidence in the synthesis step, and include a confidence score with the final answer so downstream consumers know how reliable it is.
Best for: Queries with conflicting source information, analytical reports, decision-support use cases.
Self-Hosted Tool Stack for Agentic RAG in 2026
You don’t need cloud services to build agentic RAG. Here’s what the self-hosted landscape looks like in 2026:
LLM Layer: Ollama
Ollama remains the easiest way to run local LLMs with an OpenAI-compatible API. In 2026, Ollama supports most major open-weight models including Llama 3.3, Mistral, Gemma 3, and Qwen 2.5. For agentic RAG workloads, a 7B or 8B model handles basic routing decisions, while a 70B model (if your hardware supports it) manages more complex multi-step reasoning.
Practical note: Ollama’s tool-calling support improved significantly in late 2025. You can define tools in the chat template and let the model decide when to invoke them — the core primitive for agentic RAG.
Agent Orchestration: LangGraph or LlamaIndex
LangGraph (from LangChain) gives you a graph-based framework for defining agent loops with explicit state management. It’s the most flexible option for complex multi-step retrieval pipelines. LlamaIndex offers a higher-level Agent framework that’s faster to set up for simpler use cases.
Which to choose: If you need fine-grained control over the agent loop (custom stop conditions, complex state), use LangGraph. If you want to get running quickly with built-in patterns, use LlamaIndex.
Vector Database: Qdrant or Chroma
Qdrant is the stronger choice for production workloads — it supports filtering, hybrid search, and performs well at scale. Chroma is simpler to set up and fine for experimentation or small deployments.
Embedding Model: Nomic Embed Text or BGE
Nomic’s embed-text models offer good quality with a long context window (8192 tokens), which matters when you’re working with larger chunks. BGE models from BAAI are a solid alternative with strong multilingual support.
Setting Up an Iterative Retrieval Agent: Step by Step
Here’s the practical build for the most common pattern — iterative retrieval with reflection — using Ollama, LangGraph, and Qdrant.
Prerequisites
- Ollama installed and running (
ollama serve) - Qdrant running (Docker:
docker run -p 6333:6333 qdrant/qdrant) - Python 3.10+, LangGraph, langchain-ollama, qdrant-client
Step 1: Define Retrieval Tools
from langchain_ollama import ChatOllama
from langchain_core.tools import tool
from qdrant_client import QdrantClient
qdrant = QdrantClient(host="localhost", port=6333)
collection_name = "documents"
@tool
def retrieve_docs(query: str, k: int = 5):
results = qdrant.search(
collection_name=collection_name,
query_vector=get_embedding(query),
limit=k
)
return "\n\n".join([r.payload.get("text", "") for r in results])
llm = ChatOllama(model="llama3.3", temperature=0, base_url="http://localhost:11434")
llm_with_tools = llm.bind_tools([retrieve_docs])
Step 2: Define the Agent State and Loop
from langgraph.graph import StateGraph, END
class AgentState:
question: str
retrieved_docs: list
search_queries: list
reflection: dict
answer: str | None
def should_continue(state):
if state.get("reflection", {}).get("confidence", 0) > 0.75:
return False
if len(state.get("search_queries", [])) >= 3:
return False
return True
def retrieve_and_reflect(state):
query = state["question"]
docs = retrieve_docs.invoke({"query": query, "k": 5})
critique_prompt = f"Given the question: {query}\nRetrieved:\n{docs}\nCritique: Return JSON."
import json
response = llm.invoke(critique_prompt)
try:
reflection = json.loads(response.content)
except:
reflection = {"answered": [], "missing": ["unknown"], "confidence": 0.5, "should_continue": False}
return {
"retrieved_docs": state["retrieved_docs"] + [docs],
"search_queries": state["search_queries"] + [query],
"reflection": reflection
}
def synthesize(state):
all_docs = "\n\n---\n\n".join(state["retrieved_docs"])
prompt = f"Question: {state['question']}\n\nContext:\n{all_docs}\n\nProvide a detailed answer."
answer = llm.invoke(prompt)
return {"answer": answer.content}
graph = StateGraph(AgentState)
graph.add_node("retrieve_reflect", retrieve_and_reflect)
graph.add_node("synthesize", synthesize)
graph.set_entry_point("retrieve_reflect")
graph.add_conditional_edges("retrieve_reflect", should_continue, {True: "retrieve_reflect", False: "synthesize"})
graph.add_edge("synthesize", END)
agent = graph.compile()
Step 3: Run It
result = agent.invoke({
"question": "What were the key changes in our data retention policy in 2026?",
"retrieved_docs": [],
"search_queries": []
})
print(result["answer"])
Iteration Budgets and Cost Control
Agentic RAG’s biggest practical risk is runaway cost. Each iteration burns tokens, and a poorly configured agent can loop 20+ times on a single query.
Set hard stop conditions:
| Condition | Recommended Limit |
|---|---|
| Maximum retrieval iterations | 3–5 |
| Maximum chunks per retrieval | 5–10 |
| Maximum total context tokens | 80% of model context window |
| Maximum latency budget | 60 seconds (adjust per use case) |
Monitor per-query cost in development. Most Ollama deployments don’t have native cost tracking — log the number of retrieval calls and estimated token usage per query during development so you know what each pattern costs in practice.
Route cheap queries to classic RAG. Not every question needs agentic retrieval. A simple factual lookup like “What is our refund policy?” should hit a fast classic RAG path. Use a router agent to classify query complexity upfront and route accordingly.
When Agentic RAG Is the Right Choice — and When It Is Not
Choose agentic RAG when:
- Queries are multi-hop — requiring 2+ pieces of information from different sources
- Questions are comparative or require synthesis across documents
- Source data spans multiple systems (vector store + SQL + API)
- High accuracy matters more than low latency
- You need the agent to recognize when it doesn’t know something
Stick with classic RAG when:
- Questions are simple factual lookups — one document section answers it
- Latency must be under 2 seconds
- Corpus is small and well-structured
- Cost sensitivity is high — agentic is 3–10x more expensive
- You’re hitting rate limits or running on constrained hardware
Hybrid routing is the practical answer. Most production systems in 2026 use a classifier to route queries to the appropriate retrieval strategy. Low-complexity queries go classic RAG. High-complexity queries go agentic. The routing itself can be a simple LLM call with a classification prompt.
Common Pitfalls
1. The critique loop says “I need more information” without naming what. This leads to re-retrieval with the same query. Fix: enforce a structured critique schema with specific missing-item fields.
2. No stop conditions — the agent loops forever. Fix: set hard limits on iteration count and latency from day one. Test the limits during development.
3. Mixing agentic RAG for all queries regardless of complexity. Fix: build a routing layer upfront. Not every question needs this.
4. Ignoring the token budget. Multi-step retrieval can generate 50,000+ tokens per query. Monitor this in development. If you’re on a metered LLM API, this gets expensive fast.
5. No evaluation framework. Agentic RAG is harder to test than classic RAG because the behavior is non-deterministic. Build eval queries with known answers and score your agent’s responses before going to production.
6. Weak embedding quality. If your embeddings don’t accurately represent your documents, agentic retrieval amplifies the problem — the agent is chasing bad signals. Invest in chunking strategy and embedding model selection before building the agent layer.
FAQ
Q: Can I run agentic RAG on a single VPS or does it need GPU hardware?
A: It depends on the model size. A 7B model runs reasonably on CPU for simple routing tasks, but the agent loop adds latency. For anything beyond basic iterative retrieval, a GPU (even a consumer-grade RTX 3060) makes a meaningful difference. The orchestration layer (LangGraph, Qdrant) runs fine on CPU.
Q: How does this differ from a standard LangChain RAG chain?
A: A standard LangChain RAG chain is a fixed pipeline — retrieve once, then generate. The agent loop in agentic RAG adds a reflection step where the model evaluates whether the retrieved content actually answers the question, and can decide to retrieve again. This extra control step is what separates agentic from classic.
Q: What’s the minimum viable setup to try this out?
A: Ollama + LlamaIndex’s built-in agent framework + Chroma. This gives you tool-calling agents with local LLMs in a single Python script. No Docker, no GPU required for small models. Scale up to Qdrant and LangGraph when you need production reliability.
Q: How do I evaluate whether my agentic RAG is working correctly?
A: Build an eval set of 20–50 queries with known answers. Run each query through your agent and score precision, recall, and whether the agent correctly identified gaps. Track iteration counts — if your agent is consistently hitting max iterations without reaching high confidence, your retrieval or embedding quality needs work.
Q: Can agentic RAG replace fine-tuning for domain-specific knowledge?
A: No — they’re complementary. Fine-tuning changes what the model can do. Agentic RAG changes what the model can know at inference time. Fine-tuning is better for teaching reasoning patterns and consistent formatting. Agentic RAG is better for keeping knowledge current and grounding responses in specific documents.
Key Takeaways
- Agentic RAG treats retrieval as a tool under agent control, not a fixed preprocessing step
- The five core patterns are iterative retrieval, query decomposition, hypothesis-driven retrieval, cross-corpus triangulation, and evidence-weighted synthesis
- Start with iterative retrieval — it’s the foundational pattern and handles the broadest range of cases
- Set hard iteration budgets and stop conditions from the start to control cost
- Route queries by complexity — cheap lookups should never hit the agentic path
- Enforce structured critique schemas to prevent useless re-retrieval loops
- The self-hosted stack (Ollama + LangGraph/LlamaIndex + Qdrant) is production-viable in 2026
- Invest in embedding quality before building the agent layer — agentic retrieval amplifies bad embeddings
Closing
Agentic RAG isn’t a replacement for classic RAG — it’s an addition to your retrieval toolkit. The right mental model is a spectrum: simple factual queries get simple retrieval. Complex research questions get agentic loops with multiple iterations and sources. The architectural decision isn’t “agentic or not” — it’s “which pattern does this query need.”
The self-hosted tooling matured significantly in 2025 and 2026. You can build a capable agentic RAG system today without touching a cloud API, with full control over your data, at hardware cost alone. The main remaining constraints are latency (local models are slower than hosted) and the engineering work required to tune retrieval quality and agent stop conditions for your specific corpus.
Start small. Get iterative retrieval working with your actual documents. Measure your baseline. Then layer in decomposition or cross-corpus triangulation only when your query shapes actually demand it.

