AI Usage & Credit
Every AI call is priced, recorded and counted against the workspace's credit.
Every call HyperSaaS makes to a paid AI service is recorded with what it cost, and counts against the credit of the workspace it was made for. A workspace that has used its credit can't start more AI work until the credit renews or the plan changes.
What's Counted
Each call is one AIUsage row, with a purpose:
| Purpose | When |
|---|---|
chat_reply | A reply in a chat, including each agent step |
translation | The translate endpoint |
transcription | Speech to text |
chat_title | Naming an unnamed chat after its first exchange |
chat_summary | Summarising older turns of a long chat |
query_rewrite | Rewriting a knowledge-base question before search |
query_embedding | Embedding the search queries |
rerank | Reranking search results |
document_embedding | Embedding a document's chunks at upload |
tool_call | A paid agent tool that returned an answer: web search, place lookup, weather |
So a knowledge-base question costs its reply plus the rewrite, embedding and rerank calls behind it, and all of them count.
The AIUsage Model
class AIUsage(BaseModel):
id = models.UUIDField(primary_key=True)
user = models.ForeignKey(User, null=True, on_delete=models.SET_NULL)
workspace = models.ForeignKey(Workspace, on_delete=models.CASCADE)
chat_session = models.ForeignKey(ChatSession, null=True, on_delete=models.SET_NULL)
model_used = models.CharField(max_length=100) # or the tool's name, for tool_call
purpose = models.CharField(max_length=32, choices=Purpose.choices)
input_tokens = models.PositiveIntegerField(default=0)
output_tokens = models.PositiveIntegerField(default=0)
input_cost_per_token = models.DecimalField(max_digits=15, decimal_places=10)
output_cost_per_token = models.DecimalField(max_digits=15, decimal_places=10)
total_cost = models.DecimalField(max_digits=15, decimal_places=8) # calculated on saveThe rate at the time of the call is stored on the row, so later price changes don't rewrite history.
Recording Usage
All recording goes through backend/workspaces/usage.py, and never raises: a failure is logged and the feature carries on.
A call made directly for someone records with its workspace and user given explicitly:
from backend.workspaces.usage import Purpose, record_ai_usage
record_ai_usage(
model="gpt-4.1-mini",
purpose=Purpose.TRANSLATION,
input_tokens=usage.prompt_tokens,
output_tokens=usage.completion_tokens,
workspace_id=workspace.id,
user_id=user.id,
)A call made deep inside a feature, such as reranking inside search, records against a context its caller set:
from backend.workspaces.usage import ai_usage_for
with ai_usage_for(workspace.id, user.id, chat_session.id):
results = search_documents(...) # rewrite, embedding and rerank calls record hereOutside any context, as in the evaluation commands, nothing is recorded. Embeddings, which return no token count, are counted with tiktoken's cl100k_base, the tokenizer of OpenAI's text-embedding-3 models.
Pricing
Token prices per model are in chat/pricing_config.py, keyed by model id:
PRICING_CONFIG = {
"gpt-4.1-mini": {
"input_cost_per_token": 4e-07,
"output_cost_per_token": 1.6e-06,
"max_input_tokens": 1047576,
"max_output_tokens": 32768,
},
# ... every supported model
}Paid tools have a fixed price per call, in TOOL_CALL_PRICES in workspaces/usage.py:
| Tool | Price per call | Based on |
|---|---|---|
| Web search | $0.015 | SerpApi Developer plan: $75 for 5,000 searches |
| Place lookup | $0.005 | Google Geocoding API: $5 per 1,000 requests |
| Weather | $0.0001 | WeatherAPI.com |
Change them to match your own plans. A tool call that fails (a missing key, bad input, an outage) records nothing.
Credit
A workspace's credit comes from its plan, via subscriptions.utils.workspace_credit(workspace), which returns (limit, used):
| Plan | Credit | Counted from |
|---|---|---|
| Free | $1, once | All AI usage in the owner's Free workspaces, together |
| Paid, during the trial | $1 (TRIAL_AI_CREDIT) | The start of the billing period |
| Paid | The plan's monthly credit | The start of the current credit month |
- Paid credit renews every month, on yearly plans too: each credit month starts on the monthly anniversary of the billing period's start.
- Free credit is one allowance per owner. Usage in all of an owner's Free workspaces counts against it together, so creating more workspaces doesn't create more credit.
Enforcement
subscriptions.utils.has_sufficient_buffer(user, workspace, model=...) runs before every chat message, regeneration, edit, translation and transcription, and before new documents are added. It refuses with 403 when:
- the workspace's plan doesn't include the chat's model: "The Free plan doesn't include …", or
- the credit used is within $0.01 of the limit: "Insufficient credits remaining in workspace …".
The check runs before the work, so a single reply can go a little past the limit. Each reply's length is capped at what its model can write.
Credit Usage API
GET /api/workspaces/{workspace_id}/credit-usage/{
"credit_limit": 8.0,
"total_cost_used": 2.4183,
"percentage_used": 30.23
}Any member of the workspace can read it.
Legacy Tables
UserCost and WorkspaceCost predate per-workspace credit. Nothing reads them for limits any more; credit is always computed from AIUsage.