HyperSaaS
BackendDocuments & RAG

S3 Uploads

Presigned URL upload flow for document files.

HyperSaaS uses presigned S3 URLs so file uploads go directly from the client to S3, bypassing the Django server entirely.

Upload Flow

┌──────────┐         ┌──────────┐         ┌─────┐
│  Client   │         │  Django   │         │ S3  │
└────┬─────┘         └────┬─────┘         └──┬──┘
     │                     │                   │
     │  1. POST /upload/   │                   │
     │  {filename, size}   │                   │
     │────────────────────►│                   │
     │                     │                   │
     │  2. Presigned PUT   │                   │
     │  URL + Document ID  │                   │
     │◄────────────────────│                   │
     │                     │                   │
     │  3. PUT file ───────┼──────────────────►│
     │     (direct to S3)  │                   │
     │  ◄─────────────────────────── 200 OK ───│
     │                     │                   │
     │  4. POST /confirm/  │                   │
     │────────────────────►│                   │
     │                     │── Celery task ───►│
     │  5. {task_id}       │   (ingestion)     │
     │◄────────────────────│                   │

Step 1: Request Upload URL

POST /api/workspaces/{workspace_id}/documents/upload/
{
  "filename": "product-guide.pdf",
  "content_type": "application/pdf",
  "file_size": 1234567,
  "name": "Product Guide",
  "description": "Internal product documentation",
  "knowledge_base_ids": ["uuid1"]
}

Validations:

  • Filename: required, max 500 chars
  • Extension: must be in DOCUMENT_ALLOWED_EXTENSIONS
  • File size: at least 1 byte, and within the workspace plan's largest file (10 MB on Free, up to 50 MB)
  • The plan's document count and storage, and the workspace's AI credit, aren't used up (403, naming the limit)
  • The per-user rate limit, THROTTLE_DOCUMENT_UPLOADS (30/hour), shared with web-page imports and reprocessing (429)

Step 2: Receive Presigned URL

{
  "document_id": "uuid",
  "upload_url": "https://bucket.s3.region.amazonaws.com/documents/ws-id/uuid/product-guide.pdf?X-Amz-...",
  "s3_key": "documents/ws-id/uuid/product-guide.pdf"
}

The server creates a Document record with processing_status="pending" and generates a presigned PUT URL valid for 1 hour. The URL is signed with the declared content type and file_size, so S3 only accepts a file of exactly that size. That's what makes the storage limit reliable: the size counted is the size stored.

Step 3: Upload to S3

The client uploads the file directly to S3 using the presigned URL:

await fetch(uploadUrl, {
  method: "PUT",
  body: file,
  headers: { "Content-Type": contentType },
});

Send the same Content-Type you declared in step 1; the browser sets Content-Length from the file.

Step 4: Confirm Upload

POST /api/workspaces/{workspace_id}/documents/{document_id}/confirm-upload/

This triggers the Celery ingestion task. The document status changes to "processing".

Step 5: Poll Status

GET /api/workspaces/{workspace_id}/documents/{document_id}/processing-status/

Poll until status is SUCCESS or FAILURE.

S3 Key Generation

Keys follow a predictable pattern ensuring uniqueness:

{prefix}/{workspace_id}/{uuid}/{sanitized_filename}

Example: documents/ws-abc123/doc-def456/product-guide.pdf

The UUID ensures that duplicate filenames within the same workspace don't collide.

Download URLs

Generate a temporary download URL for the original file:

GET /api/workspaces/{workspace_id}/documents/{document_id}/download-url/
{
  "download_url": "https://bucket.s3.region.amazonaws.com/documents/...?X-Amz-..."
}

Presigned GET URLs expire after 1 hour (configurable via DOCUMENT_PRESIGNED_URL_EXPIRY). They download the file as an attachment named after the original file, with a content type from its extension, whatever type the uploader declared. A file can't open as a web page from the bucket.

Deletion

However a document is deleted (through the API, or with its workspace or its owner's account):

  1. All DocumentChunk rows are cascade-deleted
  2. The Document record is removed
  3. A post_delete signal deletes the S3 object once the transaction commits. A rolled-back delete keeps the file, and an S3 error is logged without failing the delete

To remove files that no document points to, for example ones left by deletions before this signal existed:

python manage.py delete_orphaned_document_files --dry-run   # list them
python manage.py delete_orphaned_document_files

It only looks under DOCUMENT_S3_PREFIX.

Configuration

SettingDefaultDescription
AWS_ACCESS_KEY_ID—AWS credentials
AWS_SECRET_ACCESS_KEY—AWS credentials
AWS_STORAGE_BUCKET_NAME—S3 bucket name
AWS_S3_REGION_NAME—AWS region
DOCUMENT_S3_PREFIXdocumentsKey prefix in bucket
DOCUMENT_PRESIGNED_URL_EXPIRY3600URL TTL in seconds
DOCUMENT_MAX_UPLOAD_SIZE52428800Largest file any plan allows (50 MB)
DOCUMENT_ALLOWED_EXTENSIONSpdf,docx,txt,csv,md,pptx,xlsxAllowed types

If AWS credentials are not configured, the upload endpoint raises S3NotConfiguredError.

On this page