HyperSaaS
BackendDocuments & RAG

Hybrid Retrieval

Query rewriting, semantic + keyword search, Reciprocal Rank Fusion and LLM reranking.

HyperSaaS uses a hybrid retrieval strategy: the question is rewritten into standalone search queries, each is run through vector similarity search and GIN-indexed PostgreSQL full-text search, the lists are fused with Reciprocal Rank Fusion (RRF), and an LLM reranks the result. Every default here was measured on Open RAG Bench.

Search Function

def search_documents(
    query: str,
    session_id: str,
    workspace_id: str,
    top_k: int = 5,
    semantic_candidates: int = 20,
    keyword_candidates: int = 20,
    history: list[dict] | None = None,
    user=None,
) -> list[dict]:

Called by the search_knowledge_base agent tool during chat conversations. history is the recent conversation, for rewriting follow-up questions. user is the person asking: only knowledge bases they can read are searched.

The model calls made along the way (rewriting, query embedding, reranking) count against the workspace's credit.

Retrieval Flow

User question (+ recent conversation)
    │
    ▼
1. Rewrite: a standalone query + up to 2 alternative phrasings (DOCUMENT_QUERY_REWRITE=llm)
    │
    ▼
2. Resolve documents: ready documents in the chat's knowledge bases
   that the person asking can read, in this workspace
    │
    ├──────────────────────────┐
    │   for each phrasing:     │
    ▼                          ▼
3a. Semantic Search        3b. Keyword Search
    (pgvector cosine,          (PostgreSQL FTS)
     similarity ≥ 0.3)         → top 20 candidates
    → top 20 candidates
    │                          │
    └──────────┬───────────────┘
               │
               ▼
4. Reciprocal Rank Fusion over every list → 20-candidate pool
               │
               ▼
5. LLM rerank against the user's own question → top 5
               │
               ▼
6. Return results with citations

Query Rewriting

"What about the second one?" means nothing to a search engine. With DOCUMENT_QUERY_REWRITE=llm (the default), a small model (DOCUMENT_QUERY_REWRITE_MODEL, gpt-4o-mini) rewrites the question into a standalone query using the recent conversation, and adds up to DOCUMENT_QUERY_REWRITE_ALTERNATIVES (2) other phrasings. Every phrasing is searched on both legs, and all the lists are fused. If rewriting fails, the question is searched as asked.

Embeds the query and finds the closest document chunks using pgvector's cosine distance:

from pgvector.django import CosineDistance

query_embedding = embeddings.embed_query(query)

chunks = (
    DocumentChunk.objects
    .filter(document_id__in=document_ids)
    .annotate(distance=CosineDistance("embedding", query_embedding))
    .order_by("distance")
    [:semantic_candidates]
)

Distance is converted to similarity: score = 1.0 - distance. Matches below DOCUMENT_MIN_SIMILARITY (0.3) are dropped before fusion, so a question the documents don't cover finds nothing rather than the least-bad chunks.

The HNSW index (m=16, ef_construction=64, cosine ops) enables approximate nearest neighbor search — fast even with millions of chunks.

Searching part of a workspace

An HNSW index walks a fixed-size candidate list before the document filter is applied. When a chat attaches only some of a workspace's documents, which is the usual case, most candidates are thrown away and the right passages silently go missing: with a tenth of the corpus attached, recall@20 measured 16%.

DOCUMENT_HNSW_ITERATIVE_SCAN=relaxed_order (the default) uses pgvector 0.8's iterative scan to keep searching until enough rows pass the filter. Recall with a tenth of the corpus attached went from 16% to 99%, at no cost when everything is attached. It's skipped automatically on older pgvector. DOCUMENT_HNSW_EF_SEARCH sets the candidate list size (0 keeps pgvector's 40).

Uses PostgreSQL full-text search over a precomputed, GIN-indexed search_vector column — the @@ operator is index-eligible, so keyword search stays fast at any corpus size (no per-query tsvector computation):

from django.contrib.postgres.search import SearchQuery, SearchRank
from django.db.models import F

search_query = SearchQuery(query, search_type="websearch", config=FTS_LANGUAGE)

chunks = (
    DocumentChunk.objects
    .filter(
        document_id__in=document_ids,
        search_vector=search_query,   # @@ operator → GIN index
    )
    .annotate(rank=SearchRank(F("search_vector"), search_query))
    .order_by("-rank")
    [:keyword_candidates]
)

The search_vector column is populated once at ingestion (chunks are immutable), so no triggers are needed. The websearch mode supports quoted phrases and -exclusions.

FTS language (DOCUMENT_FTS_LANGUAGE)

The text-search config defaults to simple (no stemming, no stopword removal) — the safe choice for mixed-language corpora, since a single-language stemmer corrupts whichever language it doesn't match. The trade-off: natural-language questions match poorly on the keyword leg ("profitable" won't match "profitability"); the semantic leg covers exactly that gap. For single-language corpora, set DOCUMENT_FTS_LANGUAGE=english (or turkish, etc.) and re-ingest so stored vectors match the query config.

Reciprocal Rank Fusion

RRF merges the two ranked lists without a learned fusion model:

def _reciprocal_rank_fusion(
    semantic_results: list,
    keyword_results: list,
    k: int = 60,
    top_k: int = 10,
) -> list:
    scores = {}
    for rank, result in enumerate(semantic_results):
        scores[chunk_id] = 1 / (k + rank + 1)
    for rank, result in enumerate(keyword_results):
        scores[chunk_id] += 1 / (k + rank + 1)
    return sorted(scores, reverse=True)[:top_k]

The constant k=60 reduces tail-heavy bias. Chunks appearing in both lists get higher fused scores.

Example scores:

ChunkSemantic RankKeyword RankRRF Score
A#0#01/61 + 1/61 = 0.0328
B#2—1/63 = 0.0159
C—#11/62 = 0.0161

Reranking

RRF ranks by list position only — it never reads the chunk text. A second stage sends the fused candidate pool to a small LLM that reorders it by actual relevance to the question:

# documents/reranker.py
results = rerank(query, fused_pool, top_k=5)   # listwise LLM rerank
AspectBehavior
DefaultOn (DOCUMENT_RERANKER=llm); set none to turn it off
StrategyListwise rerank via DOCUMENT_RERANK_MODEL (default gpt-4o-mini), no extra packages or vendors
PoolDOCUMENT_RERANK_CANDIDATES (default 20) fused candidates
What it readsThe first DOCUMENT_RERANK_CHARS (1,500) characters of each chunk, about three quarters of a 512-token chunk
What it judges againstDOCUMENT_RERANK_QUERY=original: the user's own words, not the rewriter's paraphrase
SafetyFail-open: any reranker error returns the pre-rerank ordering — retrieval never breaks
RobustnessThe model's ranking is sanitised (invalid/duplicate indices dropped, omissions appended) so no chunk is lost
Cost~1 small-model call and about a second per search

Judging the user's own question over most of each chunk is what made reranking pay off: it took the right passage into the top five for 6.4 more questions in every hundred on the benchmark, where switching it on as it was had been within noise.

Result Format

Each result includes full citation metadata:

{
  "chunk_id": "uuid",
  "content": "The chunk text content...",
  "document_id": "uuid",
  "document_name": "Product Guide.pdf",
  "source_type": "file",
  "source_url": "",
  "chunk_index": 5,
  "page_number": 12,
  "section_heading": "Installation",
  "chunk_metadata": {},
  "score": 0.0328
}

For YouTube sources, source_url contains the video URL and chunk metadata includes timestamps.

Agent Tool Integration

The RAG tool wraps search_documents as a plain function, then each agent framework adds its own decorator:

# documents/rag_tool.py — framework-agnostic
def search_knowledge_base_impl(query: str, session, user=None) -> str:
    results = search_documents(
        query=query,
        session_id=str(session.id),
        workspace_id=str(session.workspace_id),
        history=fetch_history(session, limit=HISTORY_MESSAGES * 2),
        user=user or session.user,
    )
    return json.dumps(results)

With knowledge bases attached, the LangGraph agent's first step for each question is this search; later steps are the model's choice. Results reach the model marked as untrusted data, and it must cite each one it uses with a ref the frontend turns into a link. The tool searches as the person asking, so in a shared chat it sees only what they can read.

URL Ingest Tool

The ingest_url tool allows the agent to add new content during conversation:

def ingest_url_impl(url: str, session, user=None) -> str:
    # Auto-detects YouTube vs web_url
    # Checks the plan's document limits and the workspace's credit
    # Refuses private and internal addresses
    # Creates Document + auto-creates session KB
    # Dispatches Celery ingestion task
    return json.dumps({"status": "processing", "document_id": doc.id})

This enables users to say "read this article" or "watch this video" and have it ingested into the session's knowledge base in real time. Only the chat's owner can add links this way, only links they wrote in the conversation, and only into a knowledge base they can write to.

On this page