HyperSaaS
BackendWorkspaces & Teams

AI Usage & Credit

Every AI call is priced, recorded and counted against the workspace's credit.

Every call HyperSaaS makes to a paid AI service is recorded with what it cost, and counts against the credit of the workspace it was made for. A workspace that has used its credit can't start more AI work until the credit renews or the plan changes.

What's Counted

Each call is one AIUsage row, with a purpose:

PurposeWhen
chat_replyA reply in a chat, including each agent step
translationThe translate endpoint
transcriptionSpeech to text
chat_titleNaming an unnamed chat after its first exchange
chat_summarySummarising older turns of a long chat
query_rewriteRewriting a knowledge-base question before search
query_embeddingEmbedding the search queries
rerankReranking search results
document_embeddingEmbedding a document's chunks at upload
tool_callA paid agent tool that returned an answer: web search, place lookup, weather

So a knowledge-base question costs its reply plus the rewrite, embedding and rerank calls behind it, and all of them count.

The AIUsage Model

class AIUsage(BaseModel):
    id = models.UUIDField(primary_key=True)
    user = models.ForeignKey(User, null=True, on_delete=models.SET_NULL)
    workspace = models.ForeignKey(Workspace, on_delete=models.CASCADE)
    chat_session = models.ForeignKey(ChatSession, null=True, on_delete=models.SET_NULL)
    model_used = models.CharField(max_length=100)  # or the tool's name, for tool_call
    purpose = models.CharField(max_length=32, choices=Purpose.choices)
    input_tokens = models.PositiveIntegerField(default=0)
    output_tokens = models.PositiveIntegerField(default=0)
    input_cost_per_token = models.DecimalField(max_digits=15, decimal_places=10)
    output_cost_per_token = models.DecimalField(max_digits=15, decimal_places=10)
    total_cost = models.DecimalField(max_digits=15, decimal_places=8)  # calculated on save

The rate at the time of the call is stored on the row, so later price changes don't rewrite history.

Recording Usage

All recording goes through backend/workspaces/usage.py, and never raises: a failure is logged and the feature carries on.

A call made directly for someone records with its workspace and user given explicitly:

from backend.workspaces.usage import Purpose, record_ai_usage

record_ai_usage(
    model="gpt-4.1-mini",
    purpose=Purpose.TRANSLATION,
    input_tokens=usage.prompt_tokens,
    output_tokens=usage.completion_tokens,
    workspace_id=workspace.id,
    user_id=user.id,
)

A call made deep inside a feature, such as reranking inside search, records against a context its caller set:

from backend.workspaces.usage import ai_usage_for

with ai_usage_for(workspace.id, user.id, chat_session.id):
    results = search_documents(...)  # rewrite, embedding and rerank calls record here

Outside any context, as in the evaluation commands, nothing is recorded. Embeddings, which return no token count, are counted with tiktoken's cl100k_base, the tokenizer of OpenAI's text-embedding-3 models.

Pricing

Token prices per model are in chat/pricing_config.py, keyed by model id:

PRICING_CONFIG = {
    "gpt-4.1-mini": {
        "input_cost_per_token": 4e-07,
        "output_cost_per_token": 1.6e-06,
        "max_input_tokens": 1047576,
        "max_output_tokens": 32768,
    },
    # ... every supported model
}

Paid tools have a fixed price per call, in TOOL_CALL_PRICES in workspaces/usage.py:

ToolPrice per callBased on
Web search$0.015SerpApi Developer plan: $75 for 5,000 searches
Place lookup$0.005Google Geocoding API: $5 per 1,000 requests
Weather$0.0001WeatherAPI.com

Change them to match your own plans. A tool call that fails (a missing key, bad input, an outage) records nothing.

Credit

A workspace's credit comes from its plan, via subscriptions.utils.workspace_credit(workspace), which returns (limit, used):

PlanCreditCounted from
Free$1, onceAll AI usage in the owner's Free workspaces, together
Paid, during the trial$1 (TRIAL_AI_CREDIT)The start of the billing period
PaidThe plan's monthly creditThe start of the current credit month
  • Paid credit renews every month, on yearly plans too: each credit month starts on the monthly anniversary of the billing period's start.
  • Free credit is one allowance per owner. Usage in all of an owner's Free workspaces counts against it together, so creating more workspaces doesn't create more credit.

Enforcement

subscriptions.utils.has_sufficient_buffer(user, workspace, model=...) runs before every chat message, regeneration, edit, translation and transcription, and before new documents are added. It refuses with 403 when:

  • the workspace's plan doesn't include the chat's model: "The Free plan doesn't include …", or
  • the credit used is within $0.01 of the limit: "Insufficient credits remaining in workspace …".

The check runs before the work, so a single reply can go a little past the limit. Each reply's length is capped at what its model can write.

Credit Usage API

GET /api/workspaces/{workspace_id}/credit-usage/
{
  "credit_limit": 8.0,
  "total_cost_used": 2.4183,
  "percentage_used": 30.23
}

Any member of the workspace can read it.

Legacy Tables

UserCost and WorkspaceCost predate per-workspace credit. Nothing reads them for limits any more; credit is always computed from AIUsage.

On this page