SII-Enigma/Qwen2.5-7B-Ins-SFT-32k

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 2, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SII-Enigma/Qwen2.5-7B-Ins-SFT-32k is a 7.6 billion parameter instruction-tuned language model based on the Qwen2.5 architecture, featuring a 32k context length. Developed by SII-Enigma, this model leverages the Adaptive Multi-Guidance Policy Optimization (AMPO) framework, which enhances reasoning efficiency and performance by adaptively integrating guidance from multiple teacher models. It is specifically designed to improve learning effectiveness and achieve superior performance compared to traditional RL or SFT methods alone.

Loading preview...

Overview

SII-Enigma/Qwen2.5-7B-Ins-SFT-32k is a 7.6 billion parameter instruction-tuned model built upon the Qwen2.5 architecture, distinguished by its integration of the Adaptive Multi-Guidance Policy Optimization (AMPO) framework. AMPO is a novel approach that intelligently uses guidance from diverse teacher models, intervening only when the primary model encounters difficulties. This method aims to enhance reasoning efficiency and overall performance by optimizing how external knowledge is utilized.

Key Capabilities

  • Adaptive Multi-Guidance Replacement: Minimizes external intervention, providing guidance only upon complete on-policy failure, which preserves the model's self-discovery while boosting reasoning efficiency.
  • Comprehension-based Guidance Selection: Improves learning by guiding the model to assimilate the most comprehensible external solutions, leading to demonstrably better performance.
  • Superior Performance: Achieves enhanced performance and efficiency compared to models trained solely with Reinforcement Learning (RL) or Supervised Fine-Tuning (SFT).
  • Multi-Guidance Pool: Utilizes a pool of teacher models including AceReason-Nemotron-1.1-7B, DeepSeek-R1-Distill-Qwen-7B, OpenR1-Qwen-7B, and Qwen3-8B(thinking) to provide diverse guidance.

Good For

This model is particularly well-suited for applications requiring robust reasoning and problem-solving, where efficient and effective learning from diverse external knowledge sources is critical. Its design makes it a strong candidate for tasks that benefit from adaptive guidance and improved assimilation of complex solutions.