formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl
The formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-sft-feedback-with-proof-repair. This model was trained on the lean_selfdistill_r1_train dataset, suggesting a specialization in formal mathematics or proof-related tasks. Its training methodology indicates an iterative refinement process, likely aimed at improving reasoning and proof generation capabilities.
Loading preview...
Model Overview
This model, formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl, is an 8 billion parameter language model. It is a fine-tuned iteration of the formalmathatepfl/qwen3-sft-feedback-with-proof-repair base model, indicating a continued focus on advanced reasoning and potentially formal verification tasks.
Training Details
The model was fine-tuned using the lean_selfdistill_r1_train dataset. Key training hyperparameters include:
- Learning Rate: 2e-06
- Batch Size: 1 (train), 8 (eval)
- Optimizer: AdamW_Torch_Fused
- Scheduler: Cosine with 0.03 warmup ratio
- Epochs: 1.0
The training utilized a multi-GPU setup with 8 devices, resulting in a total effective training batch size of 8. This configuration suggests a rigorous fine-tuning process aimed at specialized performance.
Potential Use Cases
Given its lineage and training dataset name, this model is likely intended for applications requiring:
- Formal Mathematics: Assisting with mathematical proofs or formal reasoning.
- Proof Repair: Identifying and correcting errors in logical proofs.
- Automated Theorem Proving: Generating or verifying mathematical theorems within formal systems like Lean.