formalmathatepfl/classic-grpo-reasoning-sft

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The formalmathatepfl/classic-grpo-reasoning-sft is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-8b-classic-grpo. This model specializes in reasoning tasks, having been specifically trained on the lean_reasoning_sft dataset. It is designed to enhance logical deduction and problem-solving capabilities within its 32768 token context window. This fine-tuned version aims to improve performance on complex reasoning challenges.

Loading preview...

Model Overview

This model, formalmathatepfl/classic-grpo-reasoning-sft, is an 8 billion parameter language model developed by formalmathatepfl. It is a fine-tuned iteration of the formalmathatepfl/qwen3-8b-classic-grpo base model, specifically optimized for reasoning tasks.

Key Capabilities

  • Enhanced Reasoning: The model has undergone supervised fine-tuning (SFT) on the lean_reasoning_sft dataset, indicating a specialization in logical deduction and problem-solving.
  • Base Architecture: Built upon the Qwen3 architecture, providing a robust foundation for language understanding and generation.
  • Context Window: Supports a substantial context length of 32768 tokens, allowing for processing and reasoning over longer inputs.

Training Details

The model was trained with a learning rate of 1e-05, using AdamW_Torch_Fused optimizer, and a cosine learning rate scheduler with a 0.05 warmup ratio. Training involved 2 epochs across 8 devices, with a total batch size of 8. This configuration aims to imbue the model with strong reasoning abilities.

Intended Use Cases

This model is particularly well-suited for applications requiring advanced logical reasoning, mathematical problem-solving, and tasks that benefit from processing structured arguments or proofs. Its fine-tuning on a reasoning-specific dataset suggests improved performance in these domains compared to general-purpose language models.