ForSureTesterSim/Qwen2.5-R1-Minny-1.5B-v1
ForSureTesterSim/Qwen2.5-R1-Minny-1.5B-v1 is a 1.5 billion parameter Qwen2.5-based "Super-Base" model developed by ForSureTesterSim, engineered through a novel hybrid merging topology. It integrates state-of-the-art latent reasoning (GRPO), mathematical logic, and coding proficiency without catastrophic forgetting. This model is optimized for downstream alignment, serving as a highly plastic initialization checkpoint for further Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO).
Loading preview...
Model Overview
ForSureTesterSim's Qwen2.5-R1-Minny-1.5B-v1 is a 1.5 billion parameter "Super-Base" model designed for high plasticity and multi-domain proficiency. It was created using a novel hybrid merging topology that combines Model Stock (geometric anchoring) and Sens-Merging (gradient-based sensitivity scaling) to fuse advanced reasoning, mathematical logic, and coding capabilities.
Key Capabilities & Methodology
- Hybrid Merging: Utilizes a "Refined Golden Triangle" topology, anchoring on a GRPO reasoning model (
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B) and injecting pure Supervised Fine-Tuned (SFT) domain experts for math (RLinf/RLinf-math-1.5B), code (agentica-org/DeepCoder-1.5B-Preview), and reasoning (mobiuslabsgmbh/DeepSeek-R1-ReDistill-Qwen-1.5B-v1.1). - Advanced Merging Algorithm: Overcomes limitations of standard sparsification methods on GRPO models and paradigm collisions between different RL models. It employs Model Stock for geometric stability and Sens-Merging to integrate SFT logic without destroying the base model's reasoning capabilities.
- Empirical Validation: Achieves superior forward-pass cross-entropy loss across Foundation, Reasoning, and Code domains compared to raw Task Arithmetic and Model Stock, demonstrating optimal absorption of latent spaces.
Intended Use Cases
- Super-Base for Downstream Alignment: Primarily designed as an initialization checkpoint for further Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO).
- High Plasticity: Retains maximum plasticity due to mathematically shielded foundational neurons, making it highly adaptable for specific conversational formatting or agentic workflows.
- Latent Representations: Holds deep latent representations of Python/C++ coding syntax, mathematical routing, and DeepSeek
<think>traces.
Limitations
- Formatting Instability: May exhibit "Formatting Schizophrenia" due to being a merged foundation model, potentially requiring downstream instruct-tuning for stable prompt template adherence.
- Small Model Constraints: As a 1.5 billion parameter model, it is subject to typical hallucination and knowledge-retrieval limitations inherent to smaller language models.