nvidia/Nemotron-Research-Reasoning-Qwen-1.5B
The Nemotron-Research-Reasoning-Qwen-1.5B is a 1.5 billion parameter language model developed by NVIDIA, built on the Qwen architecture with a 131072 token context length. It is specifically optimized for complex reasoning tasks, including mathematical problems, coding challenges, scientific questions, and logic puzzles. Trained using the ProRL algorithm, this model significantly outperforms other 1.5B models and matches or exceeds 7B models in reasoning benchmarks, making it suitable for advanced research and development in AI reasoning.
Loading preview...
Nemotron-Research-Reasoning-Qwen-1.5B Overview
NVIDIA's Nemotron-Research-Reasoning-Qwen-1.5B is a 1.5 billion parameter model designed for advanced reasoning tasks. It leverages the Qwen architecture and boasts an extensive 131072 token context length, making it highly capable for complex problem-solving.
Key Capabilities and Differentiators
- Optimized for Reasoning: This model excels in mathematical problems, coding challenges, scientific questions, and logic puzzles, positioning it as a leading 1.5B open-weight model in these domains.
- ProRL Algorithm: Trained using the Prolonged Reinforcement Learning (ProRL) algorithm, which enables extended RL training periods and deeper exploration of reasoning strategies across diverse datasets. ProRL incorporates techniques like mitigating entropy collapse, decoupled clip and dynamic sampling policy optimization (DAPO), and KL regularization.
- Superior Performance: Demonstrates significant performance gains over DeepSeek-R1-1.5B, with average pass@1 improvements of 14.7% on math, 13.9% on coding, 54.8% on logic puzzles, 25.1% on STEM reasoning, and 18.1% on instruction-following tasks. It also matches or surpasses DeepSeek-R1-7B in various benchmarks.
- Iterative Enhancements: The model has seen subsequent versions like Nemotron-Research-Reasoning-Qwen-1.5B-v2 and Nemotron-Research-Reasoning-Qwen-1.5B-BroRL, which further scale training steps and sample sizes, leading to continuous performance improvements and setting new state-of-the-art results among 1.5B reasoning models.
Ideal Use Cases
This model is primarily intended for research and development in areas requiring strong reasoning capabilities, such as:
- Developing AI systems for complex problem-solving in STEM fields.
- Benchmarking and advancing reinforcement learning techniques for LLMs.
- Applications demanding high accuracy in logical deduction and code generation at a smaller parameter count.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.