unint64/Affine-5fudjek6tg-p28
The unint64/Affine-5fudjek6tg-p28 model is a 35.1 billion parameter Qwen3.5-MoE architecture, developed by unint64, and fine-tuned using the Reason v4 GRPO method. This model is specifically optimized for generating high-quality 'thought' sequences (z) in a mining loop context, demonstrating improved reasoning capabilities. It is designed for tasks requiring nuanced internal reasoning, particularly within coding-related conversational turns.
Loading preview...
Overview
unint64/Affine-5fudjek6tg-p28 is a 35.1 billion parameter model based on the Qwen3.5-MoE architecture, developed by unint64. It represents a candidate from the SN120 duel, fine-tuned using the Reason v4 GRPO (Gradient Regularized Policy Optimization) method. This model's lineage traces back through iterative GRPO applications, starting from a live king model and applying the same recipe to its predecessor (p24).
Key Capabilities & Training
- Reasoning Optimization: Trained with a teacher-anchored Reason v4 GRPO method, focusing on optimizing per-sample reward based on 'thought' sequences (z).
- Teacher Model: Utilizes the frozen Affine teacher
zai-org/GLM-4.5-Air-FP8for guidance during training. - Data Focus: Trained on a specific Affine public turn corpus D, comprising SWE-style coding turns, ensuring relevance to coding-related conversational contexts.
- LoRA Fine-tuning: Employs LoRA (Low-Rank Adaptation) with specific hyperparameters (r=16, alpha=128, dropout=0.05) targeting key projection modules.
- Performance: Local simulations against a previous king model showed a positive margin in mean Reason score, indicating enhanced reasoning quality, though not definitively crown-winning in live duel terms.
Use Cases
- Advanced Reasoning Tasks: Suitable for applications requiring sophisticated internal thought generation, particularly in coding or technical problem-solving scenarios.
- Mining Loop Integration: Designed for integration into mining loops where iterative refinement of reasoning capabilities is crucial.
- Code-Related Dialogue: Excels in generating coherent and relevant 'thoughts' within chat contracts that involve coding turns and bash fence structures.