Explaining how subscription tiers translate to concurrent inference call maximums.
Plans
Explaining how our different subscription tiers work.
Featherless provides serverless access to AI models and agent runtimes, allowing you to power AI applications without managing your own inference infrastructure.
Featherless offers two primary types of plans:
Feather Chat plans are subscription and concurrent-unit based, providing unlimited monthly requests with a fixed number of concurrent units. Subscription tiers vary by supported model sizes, context lengths, and the number of concurrent units included. These plans range from smaller consumer tiers designed for interactive chat, assistants, and role-playing to larger plans built for coding and agentic workloads.
Feather Developer plans are subscription and credit based. Instead of unlimited requests tied to concurrent units, usage is deducted from a monthly credit balance based on the models and inference consumed. Developer plans are designed for API-driven applications and workloads where usage may vary over time.
Chat Plans
Plan | Tier | Price (/month) | Features |
Chat | $25 |
| |
*Coming Soon | - | - | - |
*smaller models allow for higher concurrency than larger models. See more below.
Developer Plans
Developer plans are scalable, allowing users to purchase larger amounts of inference to for coding or to power production applications - whether agent fleets or other AI applications.
Plan | Price (/unit/month) | Features |
$50+ |
|
Dedicated GPU
Dedicated GPU deployments are available for workloads that need more capacity, custom infrastructure, or greater control than our standard plans provide. They are a good fit for larger models, specialized configurations, consistently high-throughput workloads, or applications whose growth has reached the point where token-based billing is no longer the most economical option.
Dedicated deployments can be tailored around your model, performance, and scaling requirements, giving you reserved capacity and predictable infrastructure costs.
GPU | RAM | GPT-OSS 120B | Gemma 4 31B | Price |
NVIDIA B300 | 288 GB | 1.5B per month | 1B per month | |
AMD MI325X | 256 GB | 500M per month | 675M per month | |
NVIDIA B200 | 180 GB | 1B per month | 700M per month | |
NVIDIA RTX PRO 6000 | 96 GB | NA | 10M per month | |
NVIDIA H100 | 80 GB | 200M per month | 70M per month |
*For more info on how the concurrent unit limits work visit: