juwon1105/RLCR-qwen3-0.6B-bigmathdigits5000

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026Architecture:Transformer Featherless Exclusive Cold

juwon1105/RLCR-qwen3-0.6B-bigmathdigits5000 is a 0.8 billion parameter language model fine-tuned from Qwen/Qwen3-0.6B. It specializes in mathematical reasoning, having been trained on the big-math-digits dataset using the GRPO method. This model is optimized for tasks requiring precise numerical and mathematical understanding, leveraging its 32768 token context length. Its primary strength lies in enhancing mathematical problem-solving capabilities within a compact model size.

Loading preview...

Overview

This model, juwon1105/RLCR-qwen3-0.6B-bigmathdigits5000, is a 0.8 billion parameter language model built upon the Qwen3-0.6B architecture. It has been specifically fine-tuned using the mehuldamani/big-math-digits dataset, focusing on improving its mathematical reasoning abilities.

Key Capabilities

  • Enhanced Mathematical Reasoning: Specialized training on a large mathematical dataset.
  • GRPO Training Method: Utilizes the GRPO (Gradient-based Reinforcement Learning with Policy Optimization) method, as introduced in the DeepSeekMath paper, to push the limits of mathematical reasoning.
  • Compact Size: At 0.8 billion parameters, it offers mathematical capabilities within a relatively small footprint.
  • TRL Framework: Trained using the TRL library for efficient fine-tuning.

Good For

  • Applications requiring strong numerical and mathematical problem-solving.
  • Research and development in mathematical reasoning with language models.
  • Scenarios where a smaller, specialized model for math tasks is preferred over larger general-purpose models.