minnesotanlp/Finch-8B-KTO
The minnesotanlp/Finch-8B-KTO is an 8 billion parameter language model based on Qwen3-8B, developed by minnesotanlp. It is fine-tuned using Evolution Fine-Tuning (EFT) and KTO preference learning, which teaches the model to evolve solutions and self-judge their quality. This model excels at mathematical discovery and competitive programming tasks, surpassing human performance on specific optimization challenges.
Loading preview...
Finch-8B-KTO: An Evolution Fine-Tuned Model for Discovery
Finch-8B-KTO is an 8 billion parameter model built upon Qwen3-8B, developed by minnesotanlp. It is distinguished by its unique two-stage training process: Evolution Fine-Tuning (EFT) followed by KTO preference learning. EFT trains the model to act as a mutation operator for evolutionary search, internalizing the process of evolving solutions. KTO then enhances this by teaching the model to self-judge the quality of candidate solutions, distinguishing between improved and regressed transitions.
Key Capabilities
- Evolutionary Search: Designed to function as a mutation operator within evolutionary scaffolds like OpenEvolve, enabling it to discover solutions for complex optimization problems.
- Self-Judgment: The KTO stage provides the model with an internal sense of solution quality, allowing it to identify promising candidates.
- Mathematical Discovery: Achieves performance that surpasses the best human scores on specific autocorrelation-inequality tasks (AC1 and AC2).
- Competitive Programming: Demonstrates significant improvements on NP-hard competitive programming tasks (e.g., FrontierCS, CALICO's P263) compared to its base model.
- Problem Solving: Outperforms its base model by +10.2% on 22 held-out tasks across 5 domains, with up to +290% improvement on individual tasks.
When to Use Finch-8B-KTO
- Optimization Tasks: Ideal for use cases requiring the discovery of novel solutions or improvements in complex optimization problems.
- Evolutionary Algorithms: Best utilized when integrated into an evolutionary scaffold (like OpenEvolve) where it can act as a mutation operator.
- Mathematical Research: Suitable for exploring and solving challenging mathematical discovery problems.
- Code Generation for Problem Solving: Can be applied to competitive programming or similar tasks where iterative improvement of code solutions is required.
Note: The model's behavior is primarily validated with the OpenEvolve scaffold; performance with other scaffolds is not guaranteed.