thu-pacman/Puro-2B-Base
Puro-2B-Base is a 2 billion parameter dense decoder-only causal language model developed by thu-pacman, pretrained from scratch on 1.4 trillion tokens. Utilizing a Qwen3-1.7B-compatible architecture with untied embeddings and blockwise FP8 training, it demonstrates competitive performance against Qwen2-1.5B and approaches Qwen2.5-1.5B on a 15-benchmark evaluation. This model is designed for research into pretraining, data recipes, and model scaling, offering an inspectable and affordable base for smaller research groups.
Loading preview...
Puro-2B-Base: An Accessible 2B-Parameter Language Model
Puro-2B-Base is a 2-billion-parameter dense causal language model developed by thu-pacman, pretrained from scratch on 1.4 trillion tokens. It features a Qwen3-1.7B-compatible architecture with untied input/output embeddings and was trained entirely on consumer-grade NVIDIA RTX 5090 GPUs, emphasizing cost-effective pretraining.
Key Capabilities & Innovations
- Cost-Efficient Pretraining: Achieves competitive performance with an estimated accelerator cost of $6,891, making large-scale pretraining more accessible.
- Novel Training Techniques: Incorporates blockwise FP8 training, the MuonH optimizer with hyperball constraints, and a two-phase data recipe.
- Strong Performance: Outperforms Qwen2-1.5B and approaches Qwen2.5-1.5B on a 15-task aggregate evaluation, particularly in math, code, reasoning, and knowledge tasks.
- Research-Oriented Release: Beyond the model weights, the project includes materialized pretraining data, training implementation, data-processing implementation, and a detailed technical report (arXiv:2608.27370).
Intended Use Cases
- Research: Ideal for studying pretraining methodologies, data recipes, optimization techniques, and model scaling.
- Base Model for Adaptation: Suitable as a compact base model for task-specific post-training and downstream adaptation.
- Inspectable Pretraining: Provides a transparent and affordable platform for smaller research groups to inspect and experiment with billion-parameter pretraining.