HyperSaaS
BackendAI Chat

API Endpoints

REST API for chat sessions, messages, and sharing.

Chat Sessions

All endpoints are scoped to a workspace, and need current membership of it (403 otherwise):

MethodEndpointDescription
GET/api/workspaces/{ws}/chats/List sessions
POST/api/workspaces/{ws}/chats/Create session
GET/api/workspaces/{ws}/chats/{id}/Retrieve session
PUT/PATCH/api/workspaces/{ws}/chats/{id}/Update session (owner only)
DELETE/api/workspaces/{ws}/chats/{id}/Delete session (owner only)
POST/api/workspaces/{ws}/chats/{id}/toggle-public/Toggle public sharing (owner only)
POST/api/workspaces/{ws}/chats/{id}/share/Share with a workspace member, by email (owner only)
GET/api/workspaces/{ws}/chats/{id}/shares/List shares
POST/api/workspaces/{ws}/chats/{id}/revoke-share/Revoke a share (owner only)

The list returns the chats you own, ones shared with you, and the workspace's public chats.

Create Session

POST /api/workspaces/{workspace_id}/chats/
{
  "name": "Research Chat",
  "ai_provider": "google",
  "ai_model": "gemini-2.5-pro",
  "system_prompt": "You are a helpful research assistant.",
  "temperature": 0.7,
  "agent_framework": "langgraph"
}

ai_provider is derived from ai_model. Leave ai_model out to get the plan's default model. Creating a chat, or switching it to another model, is refused if the workspace's plan doesn't include that model or its credit is used up.

model_parameters takes extra generation settings the provider supports, such as seed, stop, logit_bias, reasoning_effort (OpenAI), or top_k and thinking (Anthropic). See AI Model Handlers. max_tokens is capped at what the model can write in one reply.

Messages

MethodEndpointDescription
GET/api/workspaces/{ws}/chats/{session}/messages/List messages
POST/api/workspaces/{ws}/chats/{session}/messages/stream/Send a message and stream the reply
POST/api/workspaces/{ws}/chats/{session}/messages/stream/stop/Stop a streaming reply
POST/api/workspaces/{ws}/chats/{session}/messages/regenerate/stream/Stream a new reply to an earlier message
POST/api/workspaces/{ws}/chats/{session}/messages/{id}/edit/stream/Edit a message and stream a new reply
POST/api/workspaces/{ws}/chats/{session}/messages/{id}/feedback/Rate an answer
POST/api/workspaces/{ws}/chats/{session}/messages/Send a message, without streaming
POST/api/workspaces/{ws}/chats/{session}/messages/regenerate/Regenerate, without streaming
GET/api/workspaces/{ws}/chats/{session}/messages/{id}/Retrieve message
PUT/PATCH/api/workspaces/{ws}/chats/{session}/messages/{id}/Edit a message, without streaming
DELETE/api/workspaces/{ws}/chats/{session}/messages/{id}/Delete a message and its replies
GET/api/workspaces/{ws}/chats/{session}/messages/agent-task-status/{task_id}/Poll an async agent task

Anyone who can read the chat can list messages and rate answers. Sending, regenerating, editing, stopping and deleting need the chat's owner or someone it's shared with; members reading a public chat get 403. Only the chat's owner can delete replies, and a user message can be deleted by its author or the chat's owner.

The endpoints that call a model are rate limited per user: THROTTLE_CHAT_MESSAGES, 20/minute by default. Over it, they return 429.

Send a Message, Streamed

POST /api/workspaces/{workspace_id}/chats/{session_id}/messages/stream/
{
  "role": "user",
  "content": "What does the contract say about renewals?",
  "stream_id": "b3f1c2d4"
}

To attach files, send multipart form data with the same fields and up to 5 files. stream_id is your own id for the stream (letters, digits, - and _, up to 64 characters), used to stop it.

Usage-limit and validation failures return normal JSON errors before any streaming starts. Otherwise the response is text/event-stream:

event: delta
data: {"text": "The contract renews"}

event: tool
data: {"id": "call_1", "name": "search_knowledge_base", "status": "started"}

event: tool
data: {"id": "call_1", "name": "search_knowledge_base", "status": "finished"}

event: delta
data: {"text": " automatically each year [1]."}

event: done
data: {"message_id": "…", "usage": {"prompt_tokens": 2104, "completion_tokens": 86, "total_tokens": 2190}, "stopped": false}

done comes once the reply is saved, with its message id. An error event, {"detail": "Failed to get response from AI assistant. Please try again."}, ends the stream if the model fails mid-way. Handlers that can't stream return the saved reply as JSON (201) instead.

Stop a Reply

POST /api/workspaces/{ws}/chats/{session}/messages/stream/stop/
{
  "stream_id": "b3f1c2d4"
}

Returns 202. The stream ends shortly after, and the part already written is saved with stopped: true.

Regenerate and Edit

POST .../messages/regenerate/stream/      {"user_message_id": "<uuid>"}
POST .../messages/{id}/edit/stream/       {"content": "The new question"}

Both stream a new reply as another branch of the user message; earlier replies are kept. Only a message's author can edit it. Regenerating is open to the message's author and the chat's owner.

Rate an Answer

POST .../messages/{id}/feedback/
{
  "rating": "down",
  "reason": "Cited the wrong section"
}

rating is up, down, or null to clear it. Only assistant messages can be rated.

Without Streaming

POST .../messages/ takes the same body and returns the reply as JSON (201). With AGENT_ASYNC_ENABLED=True, agent work goes to a Celery task instead: the response is 202 with the user message and an agent_task_id. Poll until it's done:

GET /api/workspaces/{ws}/chats/{session}/messages/agent-task-status/{agent_task_id}/

The task must belong to a message in that chat.

Public Shared Sessions

GET /api/chats/shared/{link_uuid}/

Returns a read-only view of a public chat session. No authentication required. It shows the chat's name, model and messages, but not its system prompt, model settings, workspace, or who wrote each message.

For a chat that's shared but not public, the link works only for the owner and the people it's shared with, while they're members of its workspace. Anyone else gets 404.

Audio Transcription

POST /api/transcribe-audio/

Multipart form with audio_file, workspace_id, and language. Uses OpenAI Whisper for speech-to-text. Audio is capped at 25 MB, and requests are limited by THROTTLE_TRANSCRIPTIONS (10/minute).

Translation

POST /api/workspaces/{ws}/translate/
{
  "text": "Hello",
  "target_language": "Spanish"
}

Translates text with an AI model. Limited by THROTTLE_TRANSLATIONS (20/minute). Both transcription and translation need workspace membership and count against its credit.

On this page