Overview
AI chat system with pluggable agent frameworks and tool calling.
The chat module is the core of HyperSaaS's AI capabilities. It provides multi-model chat sessions with a pluggable agent architecture supporting LangGraph, PydanticAI, and custom frameworks.
Core Models
ChatSession
class ChatSession(BaseModel, WorkspaceAwareModel):
id = models.UUIDField(primary_key=True)
user = models.ForeignKey(User, on_delete=models.CASCADE)
workspace = models.ForeignKey(Workspace, on_delete=models.CASCADE)
name = models.CharField(max_length=255, null=True, blank=True)
status = models.CharField(choices=[ACTIVE, COMPLETED, ARCHIVED])
# AI Configuration
ai_provider = models.CharField(choices=[OPENAI, ANTHROPIC, GOOGLE, GROQ, DEEPSEEK])
ai_model = models.CharField(max_length=100)
system_prompt = models.TextField(blank=True)
temperature = models.DecimalField(null=True) # 0.0 - 2.0
max_tokens = models.PositiveIntegerField(null=True)
top_p = models.DecimalField(null=True) # 0.0 - 1.0
frequency_penalty = models.DecimalField(null=True) # -2.0 - 2.0
presence_penalty = models.DecimalField(null=True) # -2.0 - 2.0
stream = models.BooleanField(default=False)
model_parameters = models.JSONField(default=dict) # Provider-specific params
# Agent Framework
agent_framework = models.CharField(
choices=[("none", "None"), ("langgraph", "LangGraph"), ("pydantic_ai", "PydanticAI")],
default="langgraph"
)
# Sharing
is_shared = models.BooleanField(default=False)
is_public = models.BooleanField(default=False)
shareable_link_id = models.UUIDField(unique=True)Message
class Message(BaseModel):
id = models.UUIDField(primary_key=True)
session = models.ForeignKey(ChatSession, on_delete=models.CASCADE)
user = models.ForeignKey(User, null=True, blank=True)
parent_message = models.ForeignKey("self", null=True, blank=True)
sequence_number = models.PositiveIntegerField() # Auto-assigned atomically
role = models.CharField(choices=[USER, ASSISTANT, SYSTEM, TOOL])
content = models.TextField() # Max 100k chars
attachments = models.JSONField(default=list) # files sent with the message
map_data = models.JSONField(null=True, blank=True) # Location data
stopped = models.BooleanField(default=False) # the user stopped this reply mid-wayA regenerated or edited reply is another child of the same user message, so earlier answers stay as branches.
MessageFeedback
class MessageFeedback(BaseModel):
message = models.ForeignKey(Message, on_delete=models.CASCADE) # an assistant message
user = models.ForeignKey(User, on_delete=models.CASCADE)
rating = models.CharField(choices=[("up", "Up"), ("down", "Down")])
reason = models.TextField(blank=True) # up to 2,000 charactersOne rating per user per answer, set or cleared through the feedback endpoint.
The sequence_number is atomically assigned using F() expressions to prevent race conditions in concurrent message creation.
Message Processing Flow
User sends message (POST .../messages/stream/)
│
▼
Checks: can post in this chat, plan includes the model,
credit left, message fits the model, rate limit
│
▼
Save user Message (atomic sequence_number), with attachments
│
▼
Get AI handler → agent_framework check
│
├─ "langgraph" → LangGraphHandler
├─ "pydantic_ai" → PydanticAIHandler
└─ "none" → BaseChatModelHandler (no tools)
│
▼
handler.stream() → server-sent events to the client
├─ delta (reply text, as it's written)
├─ tool (a tool call starting or finishing)
└─ done (the saved reply) or error
│
▼
process_and_save_ai_response()
├─ Save assistant Message(s)
├─ Extract map_data
├─ Record AIUsage against the workspace's credit
└─ Increment session message countHandlers that can't stream fall back to invoke() and a normal JSON response. The non-streaming POST .../messages/ endpoint remains; with AGENT_ASYNC_ENABLED=True it hands agent work to a Celery task and returns 202 with an id to poll.
If the user presses Stop, the stream ends and the part already written is saved with stopped=True.
Conversation History
The model gets as much recent history as fits CHAT_HISTORY_TOKEN_BUDGET (24,000 tokens by default), counted with tiktoken, rather than a fixed number of messages. Older turns are summarised by a small model (CHAT_SUMMARY_MODEL) when CHAT_HISTORY_SUMMARIES is on, so long conversations keep their thread. A message that can't fit the model on its own, with its attachments, is refused with a clear error.
A new chat with no name is named after its first exchange (CHAT_AUTO_TITLES, CHAT_TITLE_MODEL).
Summaries and titles count against the workspace's credit like replies do.
Attachments
A message can carry up to 5 files of up to 10 MB each, sent as multipart form data:
- Images (PNG, JPEG, WebP, GIF) go to the model as vision input with that message. Only their name and size are stored, so later turns see
[Attached image: name]. Models that can't read images refuse them. - Documents of the types the knowledge base accepts are read into text, capped at 30,000 characters, and stored with the message so later turns see them too.
Attached text, search results and tool output reach the model marked as untrusted data, with rules telling it not to follow instructions found inside them.
Sharing
| Who | Read | Post, regenerate, edit, stop | Delete messages |
|---|---|---|---|
| The chat's owner | ✓ | ✓ | Any message |
| People it's shared with | ✓ | ✓ | Their own messages |
| Other workspace members, if the chat is public | ✓ | ||
| Anyone with the public link | ✓ (reduced view) |
All of these need current membership of the workspace, except the public link. The link's view leaves out the system prompt, the model settings, the workspace and who wrote each message.
Supported Providers & Models
| Provider | Models |
|---|---|
| OpenAI | gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1, gpt-4.1-mini, o3, o4-mini |
| Anthropic | claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4-5, claude-haiku-4-5 |
| gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite | |
| Which models a chat can use depends on its workspace's plan: Fast models on Free, Fast and Standard on Pro, all of them on Team and Business. |
Handlers for Groq and DeepSeek exist, but no models are offered for them. To use one, add its models to PROVIDER_MODEL_MAP in chat/model_choices.py and to a plan's tier, add pricing to chat/pricing_config.py, and define GROQ_API_KEY or DEEPSEEK_API_KEY in settings.
Tracing
With LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY set, every chat turn is recorded as a Langfuse trace, with its model calls and tool calls nested under it. Ratings are attached as scores. Set LANGFUSE_BASE_URL for a self-hosted instance or the US region.