formalmathatepfl/qwen3-sft-feedback-with-proof-repair
The formalmathatepfl/qwen3-sft-feedback-with-proof-repair is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-cpt. This model is designed for specialized tasks, leveraging its Qwen3 base architecture and a 32768 token context length. It focuses on applications derived from its fine-tuning on the sft dataset, indicating a potential for specific instruction-following or feedback-based interactions.
Loading preview...
Model Overview
This model, formalmathatepfl/qwen3-sft-feedback-with-proof-repair, is an 8 billion parameter language model. It is a fine-tuned variant of the formalmathatepfl/qwen3-cpt base model, developed by formalmathatepfl. The fine-tuning process utilized the 'sft' dataset, suggesting an optimization for specific supervised fine-tuning tasks, potentially involving feedback mechanisms or proof repair as indicated by its name.
Key Training Details
The model was trained with a learning rate of 2e-05 over 1.0 epochs, using an AdamW optimizer. It leveraged a distributed training setup across 8 GPUs, with a total train batch size of 8. The training incorporated a cosine learning rate scheduler with a 0.05 warmup ratio. The development environment included Transformers 4.57.3 and PyTorch 2.9.0+cu128.
Intended Use Cases
While specific intended uses and limitations are not detailed in the provided information, the model's fine-tuning on an 'sft' dataset implies its suitability for tasks requiring precise instruction following or generation based on structured feedback. Its Qwen3 base and 32768 token context length make it capable of handling complex and lengthy inputs, potentially in domains requiring logical reasoning or formal language processing.