// blog
RAG Tools: A Production Engineer's Stack Checklist
The engineering guide to RAG tools: chunking strategies, hybrid retrieval, faithfulness thresholds, and CI gates that catch regressions before they ship.
A RAG tool is any software primitive at one of four layers in a retrieval-augmented generation pipeline: indexing (chunking and embedding), retrieval, generation, and evaluation. Every RAG system has exactly two failure modes — retrieval failure, where wrong context reaches the model, and generation failure, where the model misrepresents good context — and the tooling that fixes each is structurally different.
What "RAG tools" actually covers
A production RAG pipeline has four layers, each with its own failure surface:
- Indexing — split documents into chunks, embed them, store in a vector index
- Retrieval — fetch the most relevant chunks at query time
- Generation — prompt the LLM with retrieved context
- Evaluation — measure whether retrieval found the right context and generation used it faithfully
Most teams invest in layer 3 (the LLM choice) and underspend on layers 1 and 4. That's backwards. Peer-reviewed research presented at NAACL 2025 (arXiv:2410.13070) found that chunking configuration had as much or more influence on retrieval quality as the choice of embedding model. Evaluation is what separates a prototype from a system you can ship on Friday.
Layer 1: Chunking — the hidden performance lever
Chunking splits source documents before embedding. The decision shapes everything downstream.
The core failure mode: a chunk reads "revenue grew by 3% over the previous quarter" but has lost the company name and date during splitting. The embedding captures a financial concept; a user query for "ACME Corp Q2 2024 revenue growth" never finds it. Correct context exists in the corpus — the split destroyed the metadata that retrieval needs.
What the benchmarks show: NVIDIA's 2024 evaluation across five datasets found that page-level chunking achieved 0.648 accuracy — highest and most consistent — for paginated documents like financial PDFs and research papers. A 2025 arXiv analysis (arXiv:2505.21700) found smaller chunks of 64–128 tokens work best for fact-dense Q&A, while 512–1024 tokens performs better for questions requiring broader context.
The safe production default: recursive character splitting at 512 tokens with 10–20% overlap. Overlap recovers the sentence that bridges two adjacent chunk boundaries.
The highest-leverage upgrade: Anthropic's Contextual Retrieval prepends a 50–100 token LLM-generated summary to each chunk before embedding and BM25 indexing, situating it in its source document. Result: a 49% reduction in retrieval failure rate, and 67% when combined with reranking — from a 5.7% baseline failure rate down to 1.9%. Cost with prompt caching: approximately $1.02 per million document tokens.
# Before Contextual Retrieval (chunk loses document context on split)
"The company's revenue grew by 3% over the previous quarter."
# After Contextual Retrieval (independently searchable)
"This chunk is from ACME Corp's Q2 2024 10-K filing.
Prior quarter revenue was $314M.
The company's revenue grew by 3% over the previous quarter."
Layer 2: Retrieval — dense, sparse, or hybrid
Dense embedding search excels at paraphrase and concept matching ("automobile" finds "car"). It fails at exact identifiers: error codes, product SKUs, API version strings, regulatory references. BM25 (sparse keyword retrieval) inverts this: strong on exact term matches, blind to synonyms and paraphrases.
Hybrid retrieval runs both in parallel and merges results with Reciprocal Rank Fusion (RRF). Dense and sparse retrieval fail in orthogonal ways — hybrid catches what either approach alone misses. RRF operates on rank positions rather than raw scores, which sidesteps the incompatibility between BM25's unbounded integers and cosine similarity in the 0–1 range. Start with RRF at k=60 as the zero-configuration default.
The reranking layer: a cross-encoder reranker scores each (query, chunk) pair together, giving a more precise relevance estimate than the first-stage retriever. Adding a cross-encoder reranker to hybrid retrieval yielded +17.2 percentage points MRR@3 and +12.1 pp Recall@5 over unreranked hybrid in a financial document benchmark. The standard pattern: retrieve top-50 with hybrid, rerank, pass top-5 to the LLM. Latency cost is typically single-digit milliseconds for a 50-candidate shortlist.
For any corpus with exact identifiers — error codes, product names, contract clauses — hybrid retrieval is not a nice-to-have. It's the correct architecture.
Layer 3: Generation — where hallucinations hide
Given retrieved context, the LLM can fail three ways:
- Hallucination: answers from parametric memory, ignoring context
- Faithful but incomplete: accurately summarizes incomplete context because retrieval undershooted
- Generation drift: stays grounded but answers a slightly different question than the one asked
Failure 2 is the hardest to catch. A faithfulness score of 1.0 combined with low context recall means the system is accurately summarizing incomplete evidence — just as dangerous as hallucination, but it reads like correct behavior. The generator is working correctly; the retriever undershooted. Fix it at layer 2, not layer 3.
System prompt grounding instructions are the cheapest first fix for failure 1, but they cannot compensate for chunks that were never retrieved.
Layer 4: Evaluation — four metrics, two failure modes
RAG evaluation separates retrieval quality from generation quality. The four standard metrics:
| Metric | What it measures | Failure type |
|---|---|---|
| Context precision | Are retrieved chunks relevant? | Retrieval noise |
| Context recall | Were all needed chunks retrieved? | Retrieval gap |
| Faithfulness | Are all claims grounded in context? | Generation hallucination |
| Answer relevancy | Does the answer address the question? | Generation drift |
Context recall is reference-based (compared against a known ideal output); the other three use an LLM-as-judge scorer and require no labeled ground truth. That makes them practical to run at scale on production traffic.
Diagnostic matrix:
- Low context recall + high faithfulness → retriever is undershooting; LLM is honest but uninformed. Fix: improve chunking, add hybrid retrieval, or apply Contextual Retrieval.
- High context precision + low faithfulness → retriever found the right chunks but LLM is hallucinating. Fix: tighten system prompt grounding instructions or upgrade the generator model.
- Low answer relevancy + high faithfulness → LLM is grounded but answering a different question than asked. Fix: revise the prompt template.
Suggested CI thresholds (calibrate to your domain):
# fn.yml — RAG pipeline eval gates
eval:
metrics:
faithfulness: ">= 0.90"
context_recall: ">= 0.80"
context_precision: ">= 0.75"
answer_relevancy: ">= 0.85"
block_on_regression: true
Wiring RAG evals into CI/CD
Evaluation that only runs before launch catches launch-day quality. Production RAG degrades as documents change, models update, and usage patterns evolve. The eval loop must run on every version.
With Functional AI, the RAG pipeline ships as a single Agentic Function. Evaluation runs against every version before promotion:
fn eval rag-pipeline@v3 \
--dataset golden-queries-v7 \
--threshold faithfulness=0.90,context-recall=0.80
# ✓ faithfulness 0.94 (+0.02 vs v2)
# ✓ context_recall 0.83 (+0.01 vs v2)
# ✓ context_precision 0.78 (+0.03 vs v2)
# ✓ answer_relevancy 0.91 (stable)
# All gates passed — v3 promoted
If a metric regresses below threshold, promotion is blocked before any user sees it. The eval dataset is versioned alongside the Agentic Function. When a production query fails, it gets added to the golden set in the same PR that fixes the retrieval gap. Trust your agents in production by measuring them before they reach it.
Who owns the eval dataset?
The ownership failure pattern: the golden query set was built in sprint one, never updated, and now scores confidently on questions no real user asks. Scores look green; production failure rate tells a different story.
Two rules that prevent this:
- The eval dataset lives in the same repo as the RAG code. Changes to chunking strategy or retrieval parameters land in the same PR as updates to the golden set.
- Every production failure is a PR. When a user reports a wrong answer, the fix includes adding that query to the golden set and updating the CI threshold gates for future deploys.
Start with 50 queries. Cover the long tail: multi-hop questions that require two separate chunks, ambiguous queries, and questions whose answer isn't in the corpus at all. Expand when production exposes a new failure mode.
FAQ
What are RAG tools? RAG tools are the software primitives at each stage of a retrieval-augmented generation pipeline: chunking libraries, embedding APIs, vector stores, hybrid retrieval setups, reranking models, and evaluation frameworks that measure context precision, recall, faithfulness, and answer relevancy.
What is the most important layer to optimize first in a RAG pipeline? Chunking and retrieval. Research shows chunking configuration influences retrieval quality as much as embedding model choice. Get recursive character splitting at 512 tokens with 10–20% overlap as your baseline, then measure context recall before tuning anything else.
When should I switch from dense-only retrieval to hybrid retrieval? As soon as your corpus contains exact identifiers — error codes, product names, version strings, regulatory references, or proper nouns. Dense-only retrieval encodes semantic intent and misses these at query time. Hybrid (BM25 + dense) with RRF at k=60 catches both failure modes and requires no score normalization.
What faithfulness score should I gate CI on for a RAG pipeline? ≥ 0.90 is a reasonable starting point for most production domains. Legal, medical, and financial applications should be higher. Set your threshold against a labeled sample of your own corpus — not benchmark data from a different domain.
How is RAG evaluation different from general LLM evaluation? General LLM evals measure output quality against a task definition. RAG evaluation separates retrieval quality (context precision and recall) from generation quality (faithfulness and answer relevancy) so you can isolate which pipeline layer is causing failures. A drop in faithfulness has a different fix than a drop in context recall.
How many golden queries do I need to start RAG evaluation? 50 is enough to start. 200+ is the target for a mature system. Diversity matters more than quantity — cover factoid questions, multi-hop questions that need two chunks, adversarial queries, and out-of-corpus questions. An eval set of 200 paraphrases of the same question is no better than one.