Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_Certainly_1p0_0p0_1p0_grpo_sapo_42_rule

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 23, 2026Architecture:Transformer Featherless Exclusive Warm

Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_Certainly_1p0_0p0_1p0_grpo_sapo_42_rule is a 2 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B-Base. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, it is optimized for tasks requiring advanced mathematical understanding and problem-solving. It is suitable for applications where robust mathematical reasoning is a primary requirement.

Loading preview...

Overview

This model, Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_Certainly_1p0_0p0_1p0_grpo_sapo_42_rule, is a specialized fine-tuned version of the Qwen3-1.7B-Base architecture, featuring approximately 2 billion parameters and a substantial context length of 32768 tokens.

Key Capabilities

  • Enhanced Mathematical Reasoning: The model was specifically trained using the GRPO (Gradient-based Reasoning Policy Optimization) method, as introduced in the DeepSeekMath paper. This training approach aims to significantly improve its performance on mathematical reasoning tasks.
  • Base Model: Built upon the robust Qwen3-1.7B-Base, providing a strong foundation for language understanding and generation.
  • Fine-tuned with TRL: The fine-tuning process leveraged the TRL (Transformers Reinforcement Learning) library, indicating a focus on optimizing specific behaviors or performance metrics.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring the model to understand, process, and generate solutions for mathematical problems.
  • Research and Development: Useful for researchers exploring the impact of GRPO on language models, particularly in the domain of mathematical reasoning.
  • Specialized Language Generation: Can be applied to tasks where the underlying mathematical reasoning capability is crucial for generating accurate and coherent responses.