juwon1105/RLCC-base-rlvr-posthoc-qwen3-1.7B-bigmathdigits5000

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026Architecture:Transformer Featherless Exclusive Cold

The juwon1105/RLCC-base-rlvr-posthoc-qwen3-1.7B-bigmathdigits5000 model is a 1.7 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B. It specializes in mathematical reasoning, having been trained on the mehuldamani/big-math-digits dataset using the GRPO method. This model is designed to enhance mathematical problem-solving capabilities, making it suitable for tasks requiring precise numerical and logical operations.

Loading preview...

Overview

This model, juwon1105/RLCC-base-rlvr-posthoc-qwen3-1.7B-bigmathdigits5000, is a specialized language model with 1.7 billion parameters, built upon the Qwen/Qwen3-1.7B architecture. Its primary focus is on advanced mathematical reasoning, achieved through fine-tuning on the mehuldamani/big-math-digits dataset.

Key Capabilities

  • Enhanced Mathematical Reasoning: Specifically trained to improve performance on tasks involving numerical and mathematical logic.
  • GRPO Training Method: Utilizes the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, for its training procedure.
  • Qwen3-1.7B Base: Leverages the robust foundation of the Qwen3-1.7B model, providing a strong general language understanding before specialization.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring accurate and complex mathematical computations or reasoning.
  • Research in Mathematical AI: Useful for researchers exploring methods to improve language models' mathematical abilities.
  • Educational Tools: Can be integrated into tools designed to assist with or generate mathematical exercises and solutions.

This model was trained using the TRL library, with specific framework versions including TRL 0.16.0.dev0, Transformers 4.51.3, Pytorch 2.5.1, Datasets 4.0.0, and Tokenizers 0.21.1.