formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-sft-feedback-with-proof-repair. This model was trained on the lean_selfdistill_r1_train dataset, suggesting a specialization in formal mathematics or proof-related tasks. Its training methodology indicates an iterative refinement process, likely aimed at improving reasoning and proof generation capabilities.

Loading preview...

Model Overview

This model, formalmathatepfl/qwen3-sft-feedback-with-proof-repair-rl, is an 8 billion parameter language model. It is a fine-tuned iteration of the formalmathatepfl/qwen3-sft-feedback-with-proof-repair base model, indicating a continued focus on advanced reasoning and potentially formal verification tasks.

Training Details

The model was fine-tuned using the lean_selfdistill_r1_train dataset. Key training hyperparameters include:

  • Learning Rate: 2e-06
  • Batch Size: 1 (train), 8 (eval)
  • Optimizer: AdamW_Torch_Fused
  • Scheduler: Cosine with 0.03 warmup ratio
  • Epochs: 1.0

The training utilized a multi-GPU setup with 8 devices, resulting in a total effective training batch size of 8. This configuration suggests a rigorous fine-tuning process aimed at specialized performance.

Potential Use Cases

Given its lineage and training dataset name, this model is likely intended for applications requiring:

  • Formal Mathematics: Assisting with mathematical proofs or formal reasoning.
  • Proof Repair: Identifying and correcting errors in logical proofs.
  • Automated Theorem Proving: Generating or verifying mathematical theorems within formal systems like Lean.