rajveer43/supply-chain-grpo-Qwen3-1.7B

Hugging Face
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 28, 2026Architecture:Transformer Featherless Exclusive Warm

The rajveer43/supply-chain-grpo-Qwen3-1.7B model is a 2 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B. It utilizes the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, for its training. This model is specialized for tasks that benefit from advanced reasoning capabilities, particularly in mathematical contexts, leveraging its 32768 token context length. Its fine-tuning with GRPO suggests an optimization for improved performance in complex problem-solving scenarios.

Loading preview...

Overview

This model, rajveer43/supply-chain-grpo-Qwen3-1.7B, is a 2 billion parameter language model derived from the Qwen3-1.7B architecture. It has been specifically fine-tuned using the TRL (Transformers Reinforcement Learning) library, incorporating a novel training procedure.

Key Capabilities & Training

The primary differentiator of this model lies in its training methodology. It was trained using GRPO (Gradient-based Reward Policy Optimization), a method detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This suggests an optimization for:

  • Enhanced reasoning abilities, particularly in mathematical or logical problem-solving.
  • Improved performance in tasks requiring structured thought processes.

Technical Details

  • Base Model: Qwen/Qwen3-1.7B
  • Parameter Count: 2 Billion
  • Context Length: 32768 tokens
  • Training Framework: TRL (Transformers Reinforcement Learning)

Potential Use Cases

Given its GRPO-based training, this model is likely well-suited for applications requiring:

  • Mathematical problem-solving and equation generation.
  • Logical deduction and complex reasoning tasks.
  • Scenarios where precise and structured outputs are critical.