Docs /Getting Started/Plans

Plans

Explaining how our different subscription tiers work.

Featherless provides serverless access to AI models and agent runtimes, allowing you to power AI applications without managing your own inference infrastructure.

Featherless offers two primary types of plans:

Feather Chat plans are subscription and concurrent-unit based, providing unlimited monthly requests with a fixed number of concurrent units. Subscription tiers vary by supported model sizes, context lengths, and the number of concurrent units included. These plans range from smaller consumer tiers designed for interactive chat, assistants, and role-playing to larger plans built for coding and agentic workloads.

Feather Developer plans are subscription and credit based. Instead of unlimited requests tied to concurrent units, usage is deducted from a monthly credit balance based on the models and inference consumed. Developer plans are designed for API-driven applications and workloads where usage may vary over time.

Chat Plans

Plan

Tier

Price (/month)

Features

Featherless Chat

Chat

$25

  • Private, secure, and anonymous usage (no logs)

  • Access any model in the catalogue (including Kimi & GLM)

  • 4 concurrent units*

  • Context size up to 32K

*Coming Soon

-

-

-

*smaller models allow for higher concurrency than larger models. See more below.

Developer Plans

Developer plans are scalable, allowing users to purchase larger amounts of inference to for coding or to power production applications - whether agent fleets or other AI applications.

Plan

Price (/unit/month)

Features

Featherless Developer

$50+

  • Credit-based API access with monthly prepaid credits

  • Pay per successful request based on model price and token usage

  • No model size limit

  • 100 concurrent units*

  • Context size up to 256K

  • See Request Pricing and Credits for billing details

Dedicated GPU

Dedicated GPU deployments are available for workloads that need more capacity, custom infrastructure, or greater control than our standard plans provide. They are a good fit for larger models, specialized configurations, consistently high-throughput workloads, or applications whose growth has reached the point where token-based billing is no longer the most economical option.

Dedicated deployments can be tailored around your model, performance, and scaling requirements, giving you reserved capacity and predictable infrastructure costs.

GPU

RAM

GPT-OSS 120B

Gemma 4 31B

Price

NVIDIA B300
288 GB · HBM3

288 GB

1.5B per month

1B per month

Contact Us

AMD MI325X
256 GB · HBM3

256 GB

500M per month

675M per month

Contact Us

NVIDIA B200
180 GB · HBM3E

180 GB

1B per month

700M per month

Contact Us

NVIDIA RTX PRO 6000
96 GB · GDDR7

96 GB

NA

10M per month

Contact Us

NVIDIA H100
80 GB · HBM2E

80 GB

200M per month

70M per month

Contact Us

*For more info on how the concurrent unit limits work visit:

Last edited: Aug 21, 2026