Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_mix_alt_Certainly_python_1p0_0p0_1p0_grpo_42_rule

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 24, 2026Architecture:Transformer Featherless Exclusive Warm

Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_mix_alt_Certainly_python_1p0_0p0_1p0_grpo_42_rule is a 2 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B-Base. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It leverages a 32768 token context length, making it suitable for tasks requiring deep contextual understanding, particularly in areas benefiting from improved reasoning.

Loading preview...

Model Overview

This model, developed by Kazuki1450, is a fine-tuned variant of the Qwen3-1.7B-Base architecture, featuring approximately 2 billion parameters and a substantial 32768 token context window. It was trained using the TRL framework.

Key Differentiator: GRPO Training

A core aspect of this model is its training methodology, which incorporates GRPO (Gradient-based Reasoning Policy Optimization). This technique, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," is specifically designed to improve a model's capabilities in mathematical reasoning. This suggests the model has been optimized to handle complex logical and numerical problems more effectively than its base counterpart.

Use Cases

Given its GRPO-enhanced training, this model is particularly well-suited for:

  • Mathematical problem-solving: Tasks requiring logical deduction and numerical computation.
  • Reasoning-intensive applications: Scenarios where robust analytical capabilities are crucial.
  • Complex question answering: Handling queries that demand more than simple information retrieval, especially those with a mathematical or logical component.

Technical Details

The model was trained with specific versions of key frameworks:

  • TRL: 0.29.0
  • Transformers: 4.57.3
  • Pytorch: 2.9.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.1