HyperSaaS
BackendDocuments & RAG

Retrieval Evaluation

Measure retrieval and answer quality, on Open RAG Bench or your own documents.

Retrieval tuning (chunk size, FTS language, fusion parameters, reranking) should be measured, not guessed. HyperSaaS ships two harnesses:

  • eval_retrieval, a quick check: a small golden question set run through each retrieval leg, reporting where the first correct chunk lands.
  • eval_open_ragbench, the full measurement: ingestion, retrieval and graded answers on Open RAG Bench, or on labelled questions over your own documents.

Results on Open RAG Bench

The whole benchmark: 1,000 arXiv papers (26,401 pages, about 55,000 chunks), all 3,045 labelled questions, and 300 answers graded against reference answers. Answers by gpt-4.1-mini through the LangGraph agent, graded by gpt-4.1, with top-k 10.

Before tuningNow
Right paper ranked first (paper hit@1)86.0%93.4%
Right paper in the top five (paper hit@5)93.0%96.2%
Right passage in the top five (section hit@5)76.0%83.4%
Correctness, 1–54.574.76
Answers grounded in retrieved passages95%97.0%
Answers that cite a source82%96.3%
Citations that point to a real search result100%99.7%
Answers saying the documents don't cover it10%3.7%

Every one of the 1,000 papers ingested, at 861 papers an hour on one 32-core machine. A full run costs a few dollars of API usage and instance time.

What changed, each measured against a control on the same questions:

  • The agent always searches first. It had answered 37% of questions from memory, and every answer graded below 4 was one of those.
  • Citing is an instruction, not an option. "Cite a result, or say you couldn't find it" made answers worse: unsure models took the exit. Keeping only the instruction raised citations without more abstentions.
  • The reranker judges the user's own question over most of each chunk. Switched on as it was, it had been within noise; this made it worth 6.4 points of section hit@5.
  • Filtered vector search keeps searching until enough rows pass the document filter. A chat attaching a tenth of a workspace had been missing most of the right passages; the benchmark attaches everything, so it never showed this.

Three changes were measured and rejected: weighting the user's words above paraphrases in fusion, removing query rewriting, and adding neighbouring chunks as context. Public benchmarks measure academic PDFs; your own documents will tell you more, and the same command runs on them.

Open RAG Bench

Three commands in documents/management/commands/:

# 1. Download: 50 papers for a pilot, or all 1,000 (about 50 minutes, at arXiv's pace)
python manage.py download_open_ragbench --pilot 50
python manage.py download_open_ragbench

# 2. Import through the real upload path: S3, then the ingestion task
python manage.py import_documents --dir ~/datasets/open_ragbench/pdfs \
  --workspace <uuid> --uploader <email> --knowledge-base "Open RAG Bench"

# 3. Evaluate
python manage.py eval_open_ragbench --knowledge-base <uuid> --ingestion-report
python manage.py eval_open_ragbench --knowledge-base <uuid> --answers 300 --workers 12 \
  --output results/full-eval.json --report results/full-eval.md
  • Ingestion report: status, parser used, timings, throughput and failures.
  • Retrieval runs the production search for every question and scores the rank of the gold paper, and of the gold section by word-sequence overlap.
  • Answers go through the real chat handler (--framework langgraph, pydantic_ai or none); a judge model grades each 1–5 against the reference and checks it's grounded. Automatic checks cover whether the knowledge base was searched, whether it cited, and whether every citation came from the search results.
  • Resumable: every finished question goes to a checkpoint, and a re-run skips what's done. Searches retry on rate limits, and failures are counted and printed above the results.

Compare two runs on the same questions:

python manage.py compare_evals before.json after.json

It counts the questions each change fixed and broke, applies McNemar's test, refuses to compare runs where more than 1% of questions failed, and won't give a verdict on slices under about 300 questions. Section scores are reported by gold-section length, because a 512-token chunk can't reach the overlap threshold against a very short section.

Your own documents: --dataset labelled takes a questions.json of questions, each with the document that answers it and, optionally, a reference answer and the passage:

[
  {
    "id": "q1",
    "query": "How long do refunds take?",
    "document": "refund-policy-2026",
    "answer": "Fourteen days from approval.",
    "section": "Refunds are issued within fourteen days of approval..."
  }
]

Fifty labelled questions over your own contracts or manuals tell you far more than any public benchmark.

The dataset is CC-BY-NC-4.0: keep it outside the repository. Parsing is CPU-bound, so for the full set use one large throwaway machine; scripts/bench_on_ec2.sh sets one up, runs everything and tears it down.

Golden-set evaluation

python manage.py eval_retrieval golden.json --workspace <uuid>
python manage.py eval_retrieval golden.json --knowledge-base <uuid> --modes keyword
python manage.py eval_retrieval golden.json --documents <uuid>,<uuid> \
    --modes hybrid,hybrid_rerank        # what reranking adds on your documents

Scope with --workspace, --knowledge-base, or --documents; select legs with --modes; control result depth with --top-k (default 10).

Golden-set format

{
  "cases": [
    {
      "description": "optional note for humans",
      "query": "What did Q3 revenue do?",
      "expect": {
        "document": "report.pdf",
        "content_contains": ["revenue grew", "gelirleri"]
      }
    }
  ]
}

A retrieved chunk matches when the optional document name substring matches and any content_contains substring appears (case-insensitive). Two deliberate design choices:

  • Content substrings, not chunk IDs — golden sets survive re-ingestion and chunking changes.
  • content_contains accepts a list — any match counts, so mixed-language corpora can accept either phrasing.

A format template ships at backend/documents/evals/golden.example.json.

Modes

ModeWhat it measures
semanticThe embedding leg alone (pgvector cosine)
keywordThe FTS leg alone — works offline, no OpenAI key needed
hybridRRF fusion, exactly as search_documents runs it
hybrid_rerankHybrid pool + LLM rerank, forced regardless of DOCUMENT_RERANKER
hybrid_rewriteHybrid with query rewriting, forced regardless of DOCUMENT_QUERY_REWRITE

Output

Per-case table (rank of the first correct chunk per mode) plus aggregates:

#   query                                     semantic   keyword    hybrid  hybrid_rerank
1   What did revenue do in the third quarter?        1      miss         1              1
2   How profitable was the company recently?         1      miss         1              1
3   hava durumu raporu                                1         1         1              1

Aggregate metrics (rank of first correct chunk):
mode            hit@1   hit@3   hit@5     MRR  evaluated
semantic         1.00    1.00    1.00   1.000          3
keyword          0.33    0.33    0.33   0.333          3
hybrid           1.00    1.00    1.00   1.000          3
hybrid_rerank    1.00    1.00    1.00   1.000          3

This sample output also illustrates the documented simple FTS trade-off: the keyword leg nails exact-phrase queries but misses natural-language questions — which the semantic leg answers at rank 1. Hybrid fusion gets the best of both.

The harness degrades gracefully: if query embedding fails (no OpenAI key), semantic/hybrid report n/a and the keyword leg still runs; if the rerank call fails, hybrid_rerank reports n/a without poisoning the other columns.

Decision workflow

  1. Ingest a real document; write ~10–20 golden cases from it.
  2. Run all modes. hybrid is your baseline.
  3. Reranking and query rewriting are on by default, because they paid off on Open RAG Bench. If hybrid_rerank doesn't beat hybrid on your documents, DOCUMENT_RERANKER=none saves a model call and about a second per search.
  4. If keyword scores near zero on a single-language corpus → set DOCUMENT_FTS_LANGUAGE to that language, re-ingest, re-run.
  5. Re-run the eval after any chunking or fusion change — the golden set is your regression suite for retrieval quality. For anything bigger, label 50+ questions and use eval_open_ragbench --dataset labelled.

On this page