Amxnn8/Qwen-2.5-3B-Reasoning-Hybrid

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 29, 2026Architecture:Transformer Featherless Exclusive Cold

Amxnn8/Qwen-2.5-3B-Reasoning-Hybrid is a 3.1 billion parameter language model, created by Amxnn8 through a SLERP merge of Qwen/Qwen2.5-3B-Instruct and Qwen/Qwen2.5-Coder-3B-Instruct. This hybrid model combines the general instruction-following capabilities of Qwen2.5-3B-Instruct with the specialized coding proficiency of Qwen2.5-Coder-3B-Instruct. It is designed to offer enhanced reasoning and coding performance within a compact 3B parameter footprint, suitable for applications requiring both general intelligence and code generation.

Loading preview...

Overview

Amxnn8/Qwen-2.5-3B-Reasoning-Hybrid is a 3.1 billion parameter language model developed by Amxnn8. It was created using the SLERP merge method, combining two distinct Qwen 2.5 models: Qwen/Qwen2.5-3B-Instruct and Qwen/Qwen2.5-Coder-3B-Instruct.

Key Capabilities

This model is engineered to leverage the strengths of its base components, aiming for a balanced performance in:

  • General instruction following: Inheriting the broad capabilities from Qwen2.5-3B-Instruct.
  • Code generation and understanding: Benefiting from the specialized training of Qwen2.5-Coder-3B-Instruct.
  • Reasoning tasks: The hybrid approach is intended to enhance overall reasoning abilities by integrating diverse knowledge domains.

Merge Details

The merge process specifically blended the layers of the two Qwen 2.5 models, with a focus on differential weighting for self-attention and MLP components. This configuration suggests an optimization strategy to balance general language understanding with coding-specific knowledge.

Good For

This model is particularly suitable for use cases that require:

  • Applications needing a compact yet capable model for both general text generation and programming tasks.
  • Scenarios where a blend of instruction-following and code-centric reasoning is beneficial.
  • Environments where resource efficiency (due to its 3B parameter size) is important without sacrificing core functionalities.