All 40,000+ open models on every plan. What changes are context, concurrency, and intended use.
For roleplay, stories and long conversations, typed by you.
For coding agents and API traffic for production applications
For teams needing a guaranteed latency fixed cost.
| Chat$25/mo | Developer$50/mo per unit | BusinessCustom | |
|---|---|---|---|
| Price & billing | |||
| Billing model | Flat rate fixed cost | Billed per token | Contract |
| Credits expire | Never | ||
| One-time top-ups | Yes | ||
| Commitment | None | None | Custom |
| Limits | |||
| Context window | 32K | up to 256K | As Requested |
| Concurrent units | 4 | 100 | Sized to your GPUs |
| Token volume | Unlimited | Metered per token | Unlimited |
| Charged for | Successful requests only | Fixed Contract Amount | |
| What you may run | |||
| Works with | SillyTavern, RisuAI, Wyvern | Cline, Kilo Code, OpenCode, OpenHands, n8n | Anything, plus private tooling |
| Interactive chat | Yes | Yes | Yes |
| Coding agents | No | Yes | Yes |
| Production API | No | Yes | Yes |
| Background automation | No | Yes | Yes |
| Reselling access | No | No | Negotiable |
| Models & capacity | |||
| Open catalogue | 40,000+ | 40,000+ | 40,000+ |
| Model size limit | None | None | None |
| Private fine-tunes | No | Yes - through hugging face | Yes |
| Capacity priority | Shared | Priority | Dedicated |
| Burst to public cloud | Yes | ||
| Support & privacy | |||
| Support | Email +Discord | Email + Discord | Named engineers |
| SLA | Yes | ||
| Onboarding | Self-serve | Self-serve | Engineer-led |
| Prompts stored | Never | Never | Never |
The chat plan is limited to human typed interactive chat use only
Character card, lorebook and history, all together.
Before the oldest turns drop out of view.
One 70B conversation, or four small ones at once.
Length is never billed. Concurrency is the only ceiling.
Your monthly credit lands in the balance each cycle. Requests draw it down at each model's own rate.
A failed request costs nothing. Credit moves only when a response comes back.
Input and output are priced separately, per million tokens. A 7B model costs a fraction of a 700B one.
Unused credit carries into the next cycle, and the one after that.
Add a one-off amount from Billing without changing your monthly credit.
Back-and-forth between you and the model is fine, whatever the frontend. What isn't covered is app or API traffic, reselling, background automation and benchmarking.
That's an HTTP 503: the model you asked for is saturated right now. Nothing is billed for the attempt. Retry or switch to another model; popular models saturate first, so a sibling model is usually free. Reach out to our support if you face this issue.
Chat boxes on model pages are for testing a model, not a full chat interface. The plan gives you an API key, which you point to SillyTavern, RisuAI, or any OpenAI-compatible frontend.
Yes, monthly. Cancel or pause from the billing page and no further charge is attempted. If the page still shows the subscription after you cancel, email support and they'll confirm it in writing.
Common, and a quick fix. Support moves you onto Developer and pro-rates what you paid, so you don't lose the month.
No. Credits stay on the org balance indefinitely. Your monthly amount tops the balance up each cycle, and you can add a one-time top-up without changing that amount.
Request-pricing calls are blocked until you add credits, and you get a notification at zero. Only successful requests are ever charged. You can top up your credits and continue.
The most common cause is too many concurrent requests, not a broken model. Developer includes 100 concurrent units; agents that fan out aggressively can exhaust that. The next most common is a model momentarily at capacity HTTP 503, never billed, switch to a different model or reach out to our support.
Tell us what you're building and we'll point you at the right plan — or tell you to stay on the cheaper one.
One OpenAI-compatible endpoint and the full catalogue on every plan. Prompts and completions are never stored.