formalmathatepfl/qwen3-sft-feedback-with-proof-repair-reasoning
The formalmathatepfl/qwen3-sft-feedback-with-proof-repair-reasoning is an 8 billion parameter Qwen3-based language model, fine-tuned from formalmathatepfl/qwen3-sft-feedback-with-proof-repair. This model specializes in reasoning tasks, particularly those involving proof repair, by leveraging feedback mechanisms. It is designed for applications requiring advanced logical deduction and correction capabilities within a 32768 token context length.
Loading preview...
Model Overview
This model, formalmathatepfl/qwen3-sft-feedback-with-proof-repair-reasoning, is an 8 billion parameter language model built upon the Qwen3 architecture. It is a fine-tuned iteration of the formalmathatepfl/qwen3-sft-feedback-with-proof-repair model, specifically trained on the lean_reasoning_sft_feedback dataset.
Key Capabilities
- Enhanced Reasoning: The model is designed to improve reasoning abilities, particularly in formal mathematical contexts.
- Proof Repair: A core focus is on the capability to identify and correct errors within proofs, leveraging a feedback-driven fine-tuning approach.
- Feedback Integration: Training incorporates feedback mechanisms to refine its logical and corrective outputs.
Training Details
The model was trained with a learning rate of 1e-05 over 2 epochs, utilizing an AdamW optimizer and a cosine learning rate scheduler with a 0.05 warmup ratio. The training involved a total batch size of 8 across 8 multi-GPU devices, ensuring robust and distributed learning.
Intended Use Cases
This model is particularly suited for applications requiring:
- Automated proof verification and correction.
- Assisted theorem proving in formal mathematics.
- Development of AI systems that can understand and repair logical arguments.