thu-pacman/Puro-2B-Base

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 16, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Puro-2B-Base is a 2 billion parameter dense decoder-only causal language model developed by thu-pacman, pretrained from scratch on 1.4 trillion tokens. Utilizing a Qwen3-1.7B-compatible architecture with untied embeddings and blockwise FP8 training, it demonstrates competitive performance against Qwen2-1.5B and approaches Qwen2.5-1.5B on a 15-benchmark evaluation. This model is designed for research into pretraining, data recipes, and model scaling, offering an inspectable and affordable base for smaller research groups.

Loading preview...

Puro-2B-Base: An Accessible 2B-Parameter Language Model

Puro-2B-Base is a 2-billion-parameter dense causal language model developed by thu-pacman, pretrained from scratch on 1.4 trillion tokens. It features a Qwen3-1.7B-compatible architecture with untied input/output embeddings and was trained entirely on consumer-grade NVIDIA RTX 5090 GPUs, emphasizing cost-effective pretraining.

Key Capabilities & Innovations

  • Cost-Efficient Pretraining: Achieves competitive performance with an estimated accelerator cost of $6,891, making large-scale pretraining more accessible.
  • Novel Training Techniques: Incorporates blockwise FP8 training, the MuonH optimizer with hyperball constraints, and a two-phase data recipe.
  • Strong Performance: Outperforms Qwen2-1.5B and approaches Qwen2.5-1.5B on a 15-task aggregate evaluation, particularly in math, code, reasoning, and knowledge tasks.
  • Research-Oriented Release: Beyond the model weights, the project includes materialized pretraining data, training implementation, data-processing implementation, and a detailed technical report (arXiv:2608.27370).

Intended Use Cases

  • Research: Ideal for studying pretraining methodologies, data recipes, optimization techniques, and model scaling.
  • Base Model for Adaptation: Suitable as a compact base model for task-specific post-training and downstream adaptation.
  • Inspectable Pretraining: Provides a transparent and affordable platform for smaller research groups to inspect and experiment with billion-parameter pretraining.