etiennebamas/feedback-grpo-step-200-reasoning-sft
The etiennebamas/feedback-grpo-step-200-reasoning-sft is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl-200-unmasked. This model specializes in reasoning tasks, specifically through fine-tuning on the lean_reasoning_sft_feedback dataset. It is designed to enhance performance in areas requiring logical deduction and structured reasoning, building upon its base Qwen3 architecture with a 32768 token context length.
Loading preview...
Model Overview
This model, etiennebamas/feedback-grpo-step-200-reasoning-sft, is an 8 billion parameter language model derived from formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl-200-unmasked. It has been specifically fine-tuned to improve its capabilities in reasoning tasks.
Key Characteristics
- Base Model: Fine-tuned from a Qwen3-based model.
- Specialization: Enhanced for reasoning tasks through supervised fine-tuning (SFT).
- Training Data: Utilizes the
lean_reasoning_sft_feedbackdataset for its specialized training. - Context Length: Supports a context window of 32768 tokens.
Training Details
The model underwent training with a learning rate of 1e-05, a total batch size of 8 across 8 GPUs, and a cosine learning rate scheduler with a 0.05 warmup ratio. The training process spanned 2 epochs, using the AdamW_TORCH_FUSED optimizer.
Potential Use Cases
This model is likely suitable for applications requiring strong logical reasoning, problem-solving, and structured output, particularly in domains where formal reasoning or proof-like structures are beneficial.