Overview
Knowledge base management with document ingestion and RAG retrieval.
The documents module provides a complete Retrieval-Augmented Generation (RAG) system — from document upload and ingestion through chunking, embedding, and hybrid search.
Core Models
KnowledgeBase
class KnowledgeBase(BaseModel, WorkspaceAwareModel):
id = models.UUIDField(primary_key=True)
name = models.CharField(max_length=200)
description = models.TextField(blank=True)
created_by = models.ForeignKey(User, on_delete=models.SET_NULL)
is_shared_with_workspace = models.BooleanField(default=False)
documents = models.ManyToManyField(Document, through="KnowledgeBaseDocument")Knowledge bases group documents together. They can be shared with an entire workspace or restricted to specific teams via KnowledgeBaseTeamAccess.
Document
class Document(BaseModel, WorkspaceAwareModel):
id = models.UUIDField(primary_key=True)
name = models.CharField(max_length=500)
uploaded_by = models.ForeignKey(User, on_delete=models.SET_NULL)
# Source
source_type = models.CharField(choices=["file", "web_url", "youtube"])
source_url = models.URLField(max_length=2048, blank=True)
# S3 Storage
s3_key = models.CharField(max_length=1024, blank=True)
original_filename = models.CharField(max_length=500, blank=True)
file_type = models.CharField(max_length=20, blank=True) # "pdf", "docx", etc.
mime_type = models.CharField(max_length=100, blank=True)
file_size = models.PositiveBigIntegerField(default=0)
# Processing
processing_status = models.CharField(
choices=["pending", "processing", "ready", "failed"]
)
processing_error = models.TextField(blank=True)
embedding_model = models.CharField(max_length=100, blank=True)
chunk_count = models.PositiveIntegerField(default=0)Documents support three source types:
- file — Uploaded via S3 presigned URL (PDF, DOCX, TXT, CSV, MD, PPTX, XLSX)
- web_url — Ingested from a web page using Trafilatura
- youtube — Transcript extracted from YouTube videos
DocumentChunk
class DocumentChunk(BaseModel):
id = models.UUIDField(primary_key=True)
document = models.ForeignKey(Document, on_delete=models.CASCADE)
chunk_index = models.PositiveIntegerField()
content = models.TextField()
# Vector embedding (pgvector)
embedding = VectorField(dimensions=1536)
embedding_model = models.CharField(max_length=100)
# Citation metadata
page_number = models.PositiveIntegerField(null=True)
section_heading = models.CharField(max_length=500, blank=True)
token_count = models.PositiveIntegerField(default=0)
chunk_metadata = models.JSONField(default=dict)
class Meta:
indexes = [
HnswIndex(fields=["embedding"], m=16, ef_construction=64, opclasses=["vector_cosine_ops"]),
models.Index(fields=["document", "chunk_index"]),
]Each chunk stores its text content, a 1536-dimensional embedding vector, and metadata for citations (page number, section heading).
Supporting Models
| Model | Purpose |
|---|---|
KnowledgeBaseDocument | Through table for KB ↔ Document M2M |
KnowledgeBaseTeamAccess | Team-level read/write access to KBs |
ChatSessionKnowledgeBase | Links chat sessions to KBs for RAG search |
DocumentProcessingTask | Tracks Celery task status for ingestion |
Processing Pipeline
Document created (pending)
│
▼
Celery task dispatched
│
├─ File → Download from S3 → Parse (Docling, or the fallback parser) → Chunk → Embed → Save
├─ Web URL → Extract (Trafilatura) → Chunk → Embed → Save
└─ YouTube → Extract transcript → Chunk → Embed → Save
│
▼
Document status → "ready" (or "failed", with a plain message for the user)A document that yields no text fails instead of being marked ready. See Ingestion Pipeline.
Retrieval Flow
User question in chat
│
▼
Rewrite into a standalone query + alternative phrasings
│
▼
Ready documents in the chat's knowledge bases that the asker can read
│
├─ Semantic search (pgvector cosine similarity) → top 20 per phrasing
├─ Keyword search (PostgreSQL full-text) → top 20 per phrasing
│
▼
Reciprocal Rank Fusion → 20-candidate pool
│
▼
LLM rerank → top 5
│
▼
Return to agent as tool result, to citeSee Hybrid Retrieval, and Retrieval Evaluation for how well it works.
Limits
What a workspace can store comes from its plan: a number of documents, total storage and a largest file size. documents.limits.check_document_allowance checks them, and the workspace's AI credit, before every upload, web-page import and link added from chat, and refuses with 403 naming the limit. Embedding a document counts against the credit.
Uploads, web-page imports and reprocessing share a per-user rate limit, THROTTLE_DOCUMENT_UPLOADS (30/hour). Reprocessing also needs credit, since it embeds the document again.
Configuration
| Setting | Default | Description |
|---|---|---|
DOCUMENT_EMBEDDING_MODEL | text-embedding-3-small | OpenAI embedding model |
DOCUMENT_EMBEDDING_DIMENSIONS | 1536 | Vector dimensions |
DOCUMENT_CHUNK_MAX_TOKENS | 512 | Max tokens per chunk |
DOCUMENT_MAX_UPLOAD_SIZE | 50MB | Largest file any plan allows |
DOCUMENT_ALLOWED_EXTENSIONS | pdf,docx,txt,csv,md,pptx,xlsx | Allowed file types |
DOCUMENT_S3_PREFIX | documents | S3 key prefix |
DOCUMENT_PRESIGNED_URL_EXPIRY | 3600 | Presigned URL TTL (seconds) |
DOCUMENT_DOCLING_PARSE_TIMEOUT | 240 | Seconds Docling may take before the fallback parser takes over |
DOCUMENT_OCR_LANGUAGE | eng | Tesseract language for the fallback parser's OCR |
DOCUMENT_RAPIDOCR_LANGUAGE | en | RapidOCR language for Docling's OCR |
Retrieval settings are listed on the Hybrid Retrieval page.
Permissions
All access goes through documents/access.py (see Permissions):
| Action | Who |
|---|---|
| See a document | Its uploader, anyone who can read a knowledge base it's in, workspace admins |
| Change, reprocess or delete a document | Its uploader, workspace admins and the owner |
| Read and search a knowledge base | Its creator, everyone if it's shared with the workspace, teams given access, workspace admins |
| Add or remove its documents | Its creator, teams given write access, workspace admins |
| Change its settings or team access, delete it | Its creator, workspace admins and the owner |