AI Model Handlers
Provider-specific LLM handlers for basic chat (no tools).
When agent_framework is set to "none", HyperSaaS falls back to basic chat using provider-specific handlers. These handle direct LLM calls without tool support.
BaseChatModelHandler
Defined in chat/ai_models.py, this base class handles:
- Message history within the token budget, with a summary of older turns (
chat.handlers.history) - Image input for models that accept it
- Token estimation with tiktoken
- Usage extraction from LLM responses
invoke(),stream()(server-sent events) and asyncastream()- Structured output via
invoke_structured()with Pydantic schemas
Provider Handlers
| Provider | Handler Class | LLM Library |
|---|---|---|
| OpenAI | OpenAIModelHandler | langchain-openai (ChatOpenAI) |
| Anthropic | AnthropicModelHandler | langchain-anthropic (ChatAnthropic) |
GoogleModelHandler | langchain-google-genai (ChatGoogleGenerativeAI) | |
| Groq | GroqModelHandler | langchain-openai with Groq endpoint |
| DeepSeek | DeepSeekModelHandler | langchain-openai with DeepSeek endpoint |
Handler Resolution
def get_ai_model_handler(session):
# 1. Check if agent framework is set
if session.agent_framework != "none":
handler = get_agent_handler(session)
if handler:
return handler
# 2. Fall back to provider-specific basic chat handler
provider = session.ai_provider
HANDLER_MAP = {
"openai": OpenAIModelHandler,
"anthropic": AnthropicModelHandler,
"google": GoogleModelHandler,
"groq": GroqModelHandler,
"deepseek": DeepSeekModelHandler,
}
return HANDLER_MAP[provider](session)Provider Configuration
Each handler reads its API key from settings:
| Provider | Setting | Env Variable |
|---|---|---|
| OpenAI | OPENAI_API_KEY | OPENAI_API_KEY |
| Anthropic | ANTHROPIC_API_KEY | ANTHROPIC_API_KEY |
GOOGLE_API_KEY | GOOGLE_API_KEY | |
| Groq | GROQ_API_KEY | GROQ_API_KEY |
| DeepSeek | DEEPSEEK_API_KEY | DEEPSEEK_API_KEY |
The Groq and DeepSeek handlers are ready, but no models are offered for them and their keys aren't defined in the base settings. See Supported Providers & Models for how to turn one on.
Session Parameters
Every chat path (LangGraph, PydanticAI and the basic handlers) reads the chat's settings through one function, generation_settings(session) in chat/model_catalog.py:
| Field | Description | When empty |
|---|---|---|
temperature | Randomness (0.0-2.0; Claude up to 1.0) | 0.2 |
max_tokens | Reply length, capped at the model's own maximum | 4,000 |
top_p | Nucleus sampling (0.0-1.0) | Provider default |
frequency_penalty | Repetition penalty (-2.0-2.0) | Provider default |
presence_penalty | Topic diversity (-2.0-2.0) | Provider default |
system_prompt | The chat's instructions, added to the system message | Empty |
model_parameters | Other generation settings the provider supports (below) | {} |
The Model Catalog
chat/model_catalog.py lists the models on offer, with each one's label, provider, tier (from the plans) and output limit, and the settings it takes. It is the one source for the model picker and its controls on the frontend (GET /api/workspaces/{id}/ai-models/) and for what each path sends, so a control the frontend shows always has an effect, and switching a chat's model never sends a setting the new model refuses.
| Models | Take |
|---|---|
| GPT-4.1, GPT-4.1 mini | Temperature, Top P, penalties, seed, stop, logit_bias, logprobs |
| GPT-5, GPT-5 mini and nano | reasoning_effort: minimal, low, medium, high |
| o3, o4-mini | reasoning_effort: low, medium, high |
| Claude | Temperature (up to 1) or Top P, not both; Top K; stop_sequences; extended thinking |
| Gemini 2.5 | Temperature, Top P, Top K |
- Reasoning models get no sampling settings. Their output limit covers their reasoning as well as the reply, so they get 16,000 tokens (
REASONING_ALLOWANCE) on top of the reply's length. - Claude gets Top P only when the chat sets it and no temperature. With extended thinking on (
{"type": "enabled", "budget_tokens": N}) it gets no temperature, Top P or Top K, and the budget comes on top of the reply's length. The budget shrinks to leave at least 1,024 tokens for the answer within the model's limit, and thinking is skipped if there isn't room for Anthropic's 1,024-token minimum. - The PydanticAI agent gets the same settings as
ModelSettings(Claude's thinking and Top K in the request body). Gemini's Top K isn't sent on that path.
max_tokens is capped at the model's max_output_tokens from chat/pricing_config.py, on every path and in the check that a message fits the model.
Model Parameters
model_parameters is free-form JSON on the chat. Only generation settings each provider accepts get through; everything else is dropped and logged. The allow-list is in chat/model_parameters.py, and the catalog then keeps only what the chat's model takes:
| Provider | Allowed keys |
|---|---|
| OpenAI, Groq, DeepSeek | seed, stop, logit_bias, logprobs, top_logprobs, reasoning_effort |
| Anthropic | top_k, stop_sequences, thinking |
top_k |
Settings that have their own field on the chat (temperature, max_tokens, top_p, the penalties) come from that field only, so each reaches the provider once. Keys that change where a request goes, what it sends or how it's billed (base_url, api_key, extra_headers, extra_body, service_tier, n) are never passed on.
{
"model_parameters": {
"seed": 7,
"reasoning_effort": "low"
}
}To offer another setting, add it to ALLOWED_MODEL_PARAMETERS and to the models' controls in the catalog. To offer another model, add it to the catalog, PRICING_CONFIG and a plan's tier.