Pricing

Same models.
Three ways to run them.

All 40,000+ open models on every plan. What changes are context, concurrency, and intended use.

Chat & roleplay
Agents & apps
Dedicated GPUs
Chat

For roleplay, stories and long conversations, typed by you.

$25/month
  • Works with SillyTavern, RisuAI, Wyvern
  • Unlimited tokens
  • 32K context
  • 4 concurrent units
  • Any model in the catalogue
Subscribe
Most chosen by builders
Developer

For coding agents and API traffic for production applications

$50/credits per month
  • Works with Cline, Pi Agent, Kilo Code, OpenCode, OpenHands, n8n and your own agents.
  • Context size up to 256K
  • 1 agent environment included
  • Fastest response time
  • Billed per token
Subscribe
Business

For teams needing a guaranteed latency fixed cost.

Custom
  • Dedicated H100 · MI325X · B200 · B300
  • No shared rate limits
  • Private fine-tunes
  • Named engineers + SLA
Talk to an engineer

Every spec, side by side

Chat$25/moDeveloper$50/mo per unitBusinessCustom
Price & billing
Billing modelFlat rate fixed costBilled per tokenContract
Credits expireNever
One-time top-upsYes
CommitmentNoneNoneCustom
Limits
Context window32Kup to 256KAs Requested
Concurrent units4100Sized to your GPUs
Token volumeUnlimitedMetered per tokenUnlimited
Charged forSuccessful requests onlyFixed Contract Amount
What you may run
Works withSillyTavern, RisuAI, WyvernCline, Kilo Code, OpenCode, OpenHands, n8nAnything, plus private tooling
Interactive chatYesYesYes
Coding agentsNoYesYes
Production APINoYesYes
Background automationNoYesYes
Reselling accessNoNoNegotiable
Models & capacity
Open catalogue40,000+40,000+40,000+
Model size limitNoneNoneNone
Private fine-tunesNoYes - through hugging faceYes
Capacity prioritySharedPriorityDedicated
Burst to public cloudYes
Support & privacy
SupportEmail +DiscordEmail + DiscordNamed engineers
SLAYes
OnboardingSelf-serveSelf-serveEngineer-led
Prompts storedNeverNeverNever
CHAT PLAN

What 32K context means on Chat Plan

The chat plan is limited to human typed interactive chat use only

~24,000
Words in play

Character card, lorebook and history, all together.

~90
Messages

Before the oldest turns drop out of view.

4
Parallel chats

One 70B conversation, or four small ones at once.

0
Tokens metered

Length is never billed. Concurrency is the only ceiling.

Developer plan

How API credits work

Your monthly credit lands in the balance each cycle. Requests draw it down at each model's own rate.

Only successes bill

A failed request costs nothing. Credit moves only when a response comes back.

Rates are per model

Input and output are priced separately, per million tokens. A 7B model costs a fraction of a 700B one.

Nothing expires

Unused credit carries into the next cycle, and the one after that.

Top up mid-cycle

Add a one-off amount from Billing without changing your monthly credit.

FAQ

Common questions

What counts as human-driven use?

Back-and-forth between you and the model is fine, whatever the frontend. What isn't covered is app or API traffic, reselling, background automation and benchmarking.

Why do I see “temporarily at capacity”?

That's an HTTP 503: the model you asked for is saturated right now. Nothing is billed for the attempt. Retry or switch to another model; popular models saturate first, so a sibling model is usually free. Reach out to our support if you face this issue.

Can I just use the chat box on the model pages?

Chat boxes on model pages are for testing a model, not a full chat interface. The plan gives you an API key, which you point to SillyTavern, RisuAI, or any OpenAI-compatible frontend.

Does it renew automatically?

Yes, monthly. Cancel or pause from the billing page and no further charge is attempted. If the page still shows the subscription after you cancel, email support and they'll confirm it in writing.

I bought Chat but I need higher context. What now?

Common, and a quick fix. Support moves you onto Developer and pro-rates what you paid, so you don't lose the month.

Do credits expire?

No. Credits stay on the org balance indefinitely. Your monthly amount tops the balance up each cycle, and you can add a one-time top-up without changing that amount.

What happens if I run out mid-run?

Request-pricing calls are blocked until you add credits, and you get a notification at zero. Only successful requests are ever charged. You can top up your credits and continue.

Why is my agent throwing errors?

The most common cause is too many concurrent requests, not a broken model. Developer includes 100 concurrent units; agents that fan out aggressively can exhaust that. The next most common is a model momentarily at capacity  HTTP 503, never billed, switch to a different model or reach out to our support.

Not sure which one?

Tell us what you're building and we'll point you at the right plan — or tell you to stay on the cheaper one.

One OpenAI-compatible endpoint and the full catalogue on every plan. Prompts and completions are never stored.