formalmathatepfl/classic-grpo-reasoning-sft
The formalmathatepfl/classic-grpo-reasoning-sft is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-8b-classic-grpo. This model specializes in reasoning tasks, having been specifically trained on the lean_reasoning_sft dataset. It is designed to enhance logical deduction and problem-solving capabilities within its 32768 token context window. This fine-tuned version aims to improve performance on complex reasoning challenges.
Loading preview...
Model Overview
This model, formalmathatepfl/classic-grpo-reasoning-sft, is an 8 billion parameter language model developed by formalmathatepfl. It is a fine-tuned iteration of the formalmathatepfl/qwen3-8b-classic-grpo base model, specifically optimized for reasoning tasks.
Key Capabilities
- Enhanced Reasoning: The model has undergone supervised fine-tuning (SFT) on the
lean_reasoning_sftdataset, indicating a specialization in logical deduction and problem-solving. - Base Architecture: Built upon the Qwen3 architecture, providing a robust foundation for language understanding and generation.
- Context Window: Supports a substantial context length of 32768 tokens, allowing for processing and reasoning over longer inputs.
Training Details
The model was trained with a learning rate of 1e-05, using AdamW_Torch_Fused optimizer, and a cosine learning rate scheduler with a 0.05 warmup ratio. Training involved 2 epochs across 8 devices, with a total batch size of 8. This configuration aims to imbue the model with strong reasoning abilities.
Intended Use Cases
This model is particularly well-suited for applications requiring advanced logical reasoning, mathematical problem-solving, and tasks that benefit from processing structured arguments or proofs. Its fine-tuning on a reasoning-specific dataset suggests improved performance in these domains compared to general-purpose language models.