Ragas vs Truvec: Evaluating RAG Retrieval Quality

A practical guide to RAG retrieval evaluation, and a comparison of Ragas and Truvec: what each measures, when to use which, and how they fit together.

  • rag
  • retrieval
  • evaluation
  • ragas
  • truvec

TL;DR: Ragas is an open-source Python library that scores a whole RAG pipeline, retrieval and generation, mostly with an LLM as judge. Truvec is a hosted tool focused on one step: retrieval. It grades embedding models and chunking on your own documents with rank-based metrics, side by side, with cost and latency. If your answers are wrong because the right passage never reached the LLM, start with Truvec. If you need to grade the answers themselves, use Ragas. Many teams need both.

I have built and debugged many RAG systems, and the most common failure I saw was not because of bad prompts or weak LLM models. Most failures were due to the right passage that was never retrieved or retrieved at the wrong rank. This post starts with a short guide on how retrieval evaluation works, then compares Ragas and Truvec and explains where each one fits.

Part 1: A short guide on RAG retrieval evaluation

What retrieval does in a RAG system

A retrieval augmented generation (RAG) system answers questions in two steps, retrieval finds the passages (chunks) of your documents that are relevant to the question and generation gives those chunks to an LLM to write the answer.

The retrieval step has many moving parts such as:

  • Chunking: How documents are split into passages (size, overlap, by paragraph, by Markdown header, by function...).
  • The embedding model: Turns each chunk and each question into a vector.
  • The retrieval algorithm: How the closest chunks are found for example cosine similarity
  • Vector storage: How the vectors are stored (Qdrant, Weaviate, PGVector...).
  • The reranker: How the closest retrieved chunks are reordered
  • Top k: How many of the chunks to pass to the LLM.

If the chunk with the answer is at rank 9 and you send the top 5, the LLM will not be able to answer the question correctly. because the generation step only works with what retrieval hands it.

What a retrieval evaluation needs

To evaluate a retrieval process, you need three things. The documents or corpus, your system will search, chunked the way production chunks them. A golden dataset, test questions linked to the chunks that contain the answer (the ground truth) and the metrics to score where the expected chunks landed in each result.

The core retrieval metrics

To explain these metrics, let's take a concrete example: Supposing you have a pdf about your company's internal processes and you want to evaluate how you can find the relevant chunks for a given question. You already chunked the pdf into 50 chunks and stored them in a vector database and created a golden dataset with questions and their expected chunks.

Take a question whose answer lives in chunks #14 and #15. The retriever configured with top_k=5, returned in order these chunks #22, #14, #8, #41, #3. This means that the first expected chunk #14 is at rank 2. The table below calculates some of the core retrieval metrics for this example:

Metric Definition Value here
Hit rate@k 1 if at least one expected chunk is in the top k 1
Recall@k Expected chunks found / expected chunks 0.5
Precision@k Expected chunks found / k 0.2
MRR 1 / rank of the first expected chunk 0.5
nDCG@k Credit for every expected chunk, discounted by rank, normalized against a perfect ranking between 0 and 1

These are deterministic retrieval metrics, the same chunks, questions and model always give the same score. Explore more in depth in The RAG metrics that actually matter.

Ground truth or LLM-as-judge (or JEV-as-judge)

There are two broad ways to decide whether retrieved chunks are relevant:

  • Reference-based or ground truth IDs: you know in advance which chunks answer each question. Scoring is exact and cheap, and you can compare models with confidence. The cost is building the golden dataset.
  • LLM-as-judge: an LLM reads the question, the retrieved chunks and the reference answer, and decides what's relevant. It needs less setup and works on any pipeline, but every score depends on a judge model, its prompt and its own mistakes, and each run costs LLM calls.

Ragas leans on the second approach. Truvec uses the first one. That difference explains most of what follows.

Part 2: Evaluation with Ragas

Ragas is an open-source Python library for evaluating LLM applications, and it's one of the most widely used tools for RAG evaluation. Using Ragas you can compute metrics such as:

  • Context precision and context recall: how relevant the retrieved contexts are, and how much of the reference they cover.
  • Faithfulness: whether the answer's claims are supported by the retrieved contexts.
  • Response relevancy: whether the answer addresses the question.
  • Other metrics, such as noise sensitivity and context entity recall.

Most of these metrics use an LLM as judge. Recent versions also include non-LLM and ID based variants of context precision and recall when you supply reference contexts. Ragas can also generate synthetic test sets from your documents.

Where you can use Ragas:

You can use Ragas to evaluate a pipeline you already built, end to end, including rerankers, hybrid search and query rewriting. and calculate metrics like faithfulness, hallucinations, relevance.

What it asks of you:

With Ragas you should build and run the pipeline for every configuration you want to test. Comparing five embedding models means indexing your corpus five times, with five provider accounts and API keys. The LLM-judged evaluations depends on the judge model and cost tokens on each run.

Part 3: Evaluation with Truvec

Truvec is a hosted platform that evaluates the retrieval step on your own documents. You upload a sample of documents (Truvec chunks them based on your preferences) or chunks that you already have, and the platform builds a golden dataset for you, runs selected embedding models against it and scores where the expected chunks land.

Where you can use Truvec:

  • Side-by-side model comparison. Pick two or more models and run one comparison. Truvec supports multiple embedding models from OpenAI, Cohere, Voyage AI, Google, and open-source models (BGE, Nomic, mxbai, GTE, Qwen3 Embedding). You don't need any provider API key: Truvec runs the models on its own accounts and your plan includes the usage.
  • Calculate rank-based metrics: hit rate, recall, precision, MRR, nDCG and F1 at your top k, plus diagnostics: hit rate at every cut-off, rank of the first relevant chunk, and the score gap between the right chunk and the best distractor.
  • Calculate cost, speed and storage: index cost, cost per 1,000 queries, query embedding latency, vector search p50/p95 and vector storage for each model, computed from real token counts.
  • The Retrieval Inspector: For every question, see the ranked chunks, marked Ground truth or Distractor, and the Expected but not retrieved chunks.
  • A Playground to run any question against any supported model, with no API setup.
  • PDF reports with a recommendation, a leaderboard, a head-to-head analysis between models and a question-by-question matrix, for the decision record.

Truvec does not evaluate the answer your LLM writes, nor does it run your own pipeline.

Truvec vs Ragas: side-by-side comparison

Ragas Truvec
Type Open-source Python library Hosted web platform with a REST API
Scope Whole RAG pipeline: retrieval and generation Retrieval only: embedding model and chunking
Relevance judgment Mostly LLM-as-judge; reference-based variants available Ground-truth chunk IDs in a reviewed, locked golden dataset
Main metrics Context precision/recall, faithfulness, response relevancy Hit rate, recall@k, precision@k, MRR, nDCG@k, F1, score gap
Comparing embedding models You index and run each model yourself +16 models in one comparison, no provider keys
Cost and latency per model Not its focus Index cost, cost per 1k queries, latency, storage
Per-question inspection Per-sample scores in a dataframe Retrieval Inspector with ranked chunks, misses and model disagreements
Test set generation Yes Yes, with review and commit workflow
Answer quality (faithfulness, hallucination) Yes No
Best for Monitoring and grading a built pipeline end to end Choosing or migrating an embedding model, tuning chunking

When to use which

Use Truvec when:

  • You're choosing an embedding model for a new project and want evidence on your own documents, not a leaderboard rank.
  • You're considering a migration to a new embedding model and need to know if it's better enough to justify re-indexing everything.
  • Your answers are often wrong or vague and you suspect retrieval. The Misses filter usually tells you in ten minutes whether the problem is the model, the chunking, PDF extraction or the test questions.
  • You want to test a different chunking setup quickly without rebuilding your pipeline.
  • You need a report a non-specialist can read.

Use Ragas when:

  • Retrieval is solid and you need to grade the generated answers: faithfulness, hallucinations, relevance.
  • You want to score the exact pipeline you run, including rerankers and query rewriting, inside your codebase.

Use both when you're building a serious RAG product. A workflow that works well:

  1. Prototype retrieval in Truvec: pick the embedding model and chunking with a golden dataset and a comparison.
  2. Build the pipeline with the winning setup.
  3. Evaluate end to end with Ragas, focusing on the generation metrics.
  4. When a new embedding model comes out or your corpus changes, go back to Truvec and re-run the comparison with your current model as the baseline.

Why isolating retrieval matters

When a RAG pipeline quality drops, you don't know whether the retriever, the prompt or the LLM is to blame. Isolating retrieval metrics like hit rate@5 helps you pinpoint the issue: if hit rate@5 is 62%, then for 38% of questions the LLM never had a chance.

It also makes decisions cheaper. Comparing six embedding models with an LLM judge means six indexing pipelines and thousands of judge calls. Comparing them on a locked golden dataset is deterministic.

FAQ

Is Truvec a Ragas alternative?

For retrieval evaluation and embedding model selection, yes. For grading generated answers (faithfulness, answer relevancy), no: Truvec doesn't evaluate the LLM's answer. The two tools cover different steps and work well together.

Does Truvec use an LLM as judge?

No. Truvec uses an AI model to generate draft questions for the golden dataset, which you then review and commit. Scoring itself compares retrieved chunk IDs with the expected ones, so it's deterministic.

Do I need API keys for OpenAI, Cohere or Voyage to compare models in Truvec?

No. Truvec runs the supported models on its own accounts, and evaluation usage is included in your plan.

Can Truvec evaluate the chunks my production pipeline already produces?

Yes. Upload them as pre-chunked data (JSON, JSONL or CSV) and Truvec evaluates exactly those chunks.

Try it on your documents

Truvec has a Free plan with no card needed: one project with up to 5 documents, a golden set of up to 20 questions, and one comparison of up to 3 commercial models. It's enough to see what your retrieval misses. Start at app.truvec.dev, and read the quickstart for a 15-minute walkthrough.