HyperSaaS
BackendAI Chat

AI Model Handlers

Provider-specific LLM handlers for basic chat (no tools).

When agent_framework is set to "none", HyperSaaS falls back to basic chat using provider-specific handlers. These handle direct LLM calls without tool support.

BaseChatModelHandler

Defined in chat/ai_models.py, this base class handles:

  • Message history within the token budget, with a summary of older turns (chat.handlers.history)
  • Image input for models that accept it
  • Token estimation with tiktoken
  • Usage extraction from LLM responses
  • invoke(), stream() (server-sent events) and async astream()
  • Structured output via invoke_structured() with Pydantic schemas

Provider Handlers

ProviderHandler ClassLLM Library
OpenAIOpenAIModelHandlerlangchain-openai (ChatOpenAI)
AnthropicAnthropicModelHandlerlangchain-anthropic (ChatAnthropic)
GoogleGoogleModelHandlerlangchain-google-genai (ChatGoogleGenerativeAI)
GroqGroqModelHandlerlangchain-openai with Groq endpoint
DeepSeekDeepSeekModelHandlerlangchain-openai with DeepSeek endpoint

Handler Resolution

def get_ai_model_handler(session):
    # 1. Check if agent framework is set
    if session.agent_framework != "none":
        handler = get_agent_handler(session)
        if handler:
            return handler

    # 2. Fall back to provider-specific basic chat handler
    provider = session.ai_provider
    HANDLER_MAP = {
        "openai": OpenAIModelHandler,
        "anthropic": AnthropicModelHandler,
        "google": GoogleModelHandler,
        "groq": GroqModelHandler,
        "deepseek": DeepSeekModelHandler,
    }
    return HANDLER_MAP[provider](session)

Provider Configuration

Each handler reads its API key from settings:

ProviderSettingEnv Variable
OpenAIOPENAI_API_KEYOPENAI_API_KEY
AnthropicANTHROPIC_API_KEYANTHROPIC_API_KEY
GoogleGOOGLE_API_KEYGOOGLE_API_KEY
GroqGROQ_API_KEYGROQ_API_KEY
DeepSeekDEEPSEEK_API_KEYDEEPSEEK_API_KEY

The Groq and DeepSeek handlers are ready, but no models are offered for them and their keys aren't defined in the base settings. See Supported Providers & Models for how to turn one on.

Session Parameters

Every chat path (LangGraph, PydanticAI and the basic handlers) reads the chat's settings through one function, generation_settings(session) in chat/model_catalog.py:

FieldDescriptionWhen empty
temperatureRandomness (0.0-2.0; Claude up to 1.0)0.2
max_tokensReply length, capped at the model's own maximum4,000
top_pNucleus sampling (0.0-1.0)Provider default
frequency_penaltyRepetition penalty (-2.0-2.0)Provider default
presence_penaltyTopic diversity (-2.0-2.0)Provider default
system_promptThe chat's instructions, added to the system messageEmpty
model_parametersOther generation settings the provider supports (below){}

The Model Catalog

chat/model_catalog.py lists the models on offer, with each one's label, provider, tier (from the plans) and output limit, and the settings it takes. It is the one source for the model picker and its controls on the frontend (GET /api/workspaces/{id}/ai-models/) and for what each path sends, so a control the frontend shows always has an effect, and switching a chat's model never sends a setting the new model refuses.

ModelsTake
GPT-4.1, GPT-4.1 miniTemperature, Top P, penalties, seed, stop, logit_bias, logprobs
GPT-5, GPT-5 mini and nanoreasoning_effort: minimal, low, medium, high
o3, o4-minireasoning_effort: low, medium, high
ClaudeTemperature (up to 1) or Top P, not both; Top K; stop_sequences; extended thinking
Gemini 2.5Temperature, Top P, Top K
  • Reasoning models get no sampling settings. Their output limit covers their reasoning as well as the reply, so they get 16,000 tokens (REASONING_ALLOWANCE) on top of the reply's length.
  • Claude gets Top P only when the chat sets it and no temperature. With extended thinking on ({"type": "enabled", "budget_tokens": N}) it gets no temperature, Top P or Top K, and the budget comes on top of the reply's length. The budget shrinks to leave at least 1,024 tokens for the answer within the model's limit, and thinking is skipped if there isn't room for Anthropic's 1,024-token minimum.
  • The PydanticAI agent gets the same settings as ModelSettings (Claude's thinking and Top K in the request body). Gemini's Top K isn't sent on that path.

max_tokens is capped at the model's max_output_tokens from chat/pricing_config.py, on every path and in the check that a message fits the model.

Model Parameters

model_parameters is free-form JSON on the chat. Only generation settings each provider accepts get through; everything else is dropped and logged. The allow-list is in chat/model_parameters.py, and the catalog then keeps only what the chat's model takes:

ProviderAllowed keys
OpenAI, Groq, DeepSeekseed, stop, logit_bias, logprobs, top_logprobs, reasoning_effort
Anthropictop_k, stop_sequences, thinking
Googletop_k

Settings that have their own field on the chat (temperature, max_tokens, top_p, the penalties) come from that field only, so each reaches the provider once. Keys that change where a request goes, what it sends or how it's billed (base_url, api_key, extra_headers, extra_body, service_tier, n) are never passed on.

{
  "model_parameters": {
    "seed": 7,
    "reasoning_effort": "low"
  }
}

To offer another setting, add it to ALLOWED_MODEL_PARAMETERS and to the models' controls in the catalog. To offer another model, add it to the catalog, PRICING_CONFIG and a plan's tier.

On this page