What's New in HyperSaaS: Plans, Metered AI, Streaming Chat and Measured RAG
Plans defined in code and synced to Stripe, every AI call counted against credit, streaming chat with attachments, and a knowledge base benchmarked on 3,045 questions.
The last few months turned HyperSaaS from a foundation into something closer to a finished product. Billing now comes with real plans and limits, every AI call is metered, chat streams, and the knowledge base has numbers behind it. Here's what changed.
Plans that ship with the code
HyperSaaS now includes four plans, defined in one file (subscriptions/plans.py) and enforced throughout the app:
| | Free | Pro | Team | Business | |---|---|---|---|---| | Price per workspace | $0 | $20/mo | $60/mo | $200/mo | | AI credit | $1, once | $8/month | $24/month | $80/month | | Documents | 25 | 500 | 2,500 | 10,000 | | Storage | 50 MB | 2 GB | 10 GB | 50 GB | | Models | Fast | Fast, Standard | All | All |
Yearly prices cost ten months. Paid plans start with a 7-day trial.
You don't create products in the Stripe dashboard any more. Edit the plans file, then run:
python manage.py sync_stripe_plans --dry-run
python manage.py sync_stripe_plans
The command creates or updates the products and prices and archives anything that isn't a plan. It also configures the Stripe billing portal so customers can switch between exactly these plans. Running it twice changes nothing.
What a plan allows (credit, models, documents, storage, file size) is always read from your code, never from Stripe. A limit change takes effect when you deploy.
Every AI call is metered
Credit is per workspace, and it now covers everything that costs you money, not only chat replies:
- chat replies, translation and transcription;
- chat titles and summaries of long conversations;
- document embeddings at upload;
- query rewriting, query embedding and reranking on every knowledge-base search;
- paid agent tools (web search, place lookups, weather), at a fixed price per call.
Each call is stored as an AIUsage row with its model, tokens, cost and purpose, so you can see where the money goes. Paid credit renews monthly, on yearly plans too.
Free credit is one allowance per account owner, shared across their Free workspaces, so creating more workspaces doesn't create more credit. When a workspace runs out, chat stops and new documents are refused, with a clear message.
Chat that streams
Replies now stream token by token over server-sent events, and the chat has caught up with what people expect:
- Stop a reply mid-way. The part already written is kept.
- Edit a message or regenerate a reply. Both stream, and earlier answers stay as branches.
- Attach up to five files or images to a message. Images go to vision models; documents are read into the conversation.
- Rate answers with thumbs up or down.
- Long conversations are summarised automatically, so history fits the model without losing the thread.
- Unnamed chats are named after their first exchange.
Each plan has its own set of models, from fast and cheap to premium. A chat's settings can include any generation setting its provider supports, and reply length is capped at what the model can actually write. Every chat turn can be traced in Langfuse, and endpoints that call a model are rate limited per user.
A knowledge base with numbers behind it
We benchmarked the whole RAG pipeline on Open RAG Bench: 1,000 arXiv papers (26,401 pages), 3,045 labelled questions and 300 answers graded against reference answers.
| | Before tuning | Now | |---|---|---| | Right paper ranked first | 86.0% | 93.4% | | Right passage in the top five | 76.0% | 83.4% | | Answers grounded in retrieved text | 95% | 97.0% | | Answers that cite a source | 82% | 96.3% | | Citations that point to a real result | 100% | 99.7% |
Every change was measured against a control on the same questions. These shipped:
- The agent always searches first. It had been answering about a third of questions from memory, and those were its worst answers.
- Citing is no longer optional. Answers that use a document cite it.
- Reranking is on by default. It now judges the user's own question against more of each passage.
- Search works in partial knowledge bases. When a chat attached only part of a workspace's documents, the vector index was silently missing most of the right passages. It now keeps searching until enough of them match, at no cost when everything is attached.
Ingestion got sturdier too. All 1,000 papers ingested, with none failing, at about 860 papers an hour on one 32-core machine. Docling runs in a process that can be stopped, with a fallback parser. OCR only runs on pages without a text layer. Only failures that clear up on their own, like network errors and rate limits, are retried.
The download, bulk-import and evaluation commands ship with HyperSaaS, so you can run the same measurement on your own documents. That tells you far more than any public benchmark.
Workspaces and accounts
- Invitations are the way into a workspace or team. Workspace owners and admins can invite to the workspace or any of its teams; team owners and admins can invite to their own team.
- Ownership transfer lets an owner hand a workspace to another member, who becomes owner and admin.
- Deleting an account first asks you to cancel your subscription and transfer any workspace other people use, so nobody loses their work.
- Changing your email sends a link to the new address. The change happens when that link is opened, and the old address is told.
- Sessions use one-hour access tokens that the frontend refreshes automatically. Each refresh issues a new refresh token.
A newer stack
The frontend is now on Next.js 16, with its shadcn/ui components moved to Base UI, along with Zod 4 and the latest AI SDK.
Upgrading
If you're running an earlier version:
- Run
python manage.py migrate. - Run
sync_stripe_plans --dry-run, then without--dry-run, in test mode first and in live mode before launch. It archives products from before plans existed; their subscribers count as Free until they move to a plan. - Remove
STRIPE_ENDPOINT_SECRETandSTRIPE_PRODUCT_IDfrom your environment; nothing reads them. - Run
delete_orphaned_document_files --dry-run, then without the flag, to remove stored files that no document points to any more.
The documentation covers each of these in detail.