dracko14/Myth-4B
dracko14/Myth-4B is a 4.5 billion parameter language model created by dracko14, formed by SLERP fusion of Qwen 3.5 4B and Qwen3.5-4B-Neo. This model combines Qwen 3.5's multimodal capabilities, coding, and math foundation with Neo's deep reasoning, offering a 32K context length. It is designed for enhanced reasoning, multimodal tasks, and strong performance in coding and mathematics.
Loading preview...
Myth 4B: A Fused Language Model for Enhanced Reasoning and Multimodality
Myth 4B is a 4.5 billion parameter model developed by dracko14, created through a unique SLERP (Spherical Linear Interpolation) fusion of two base models: Qwen 3.5 4B and Qwen3.5-4B-Neo. This fusion process was executed using a custom engine that operates directly on safetensors, bypassing traditional mergekit requirements and allowing for low-memory merging.
Key Capabilities & Features
- Deep Reasoning: Inherits advanced reasoning capabilities from the Neo model's specialized fine-tuning.
- Multimodal Support: Retains the native multimodal architecture of Qwen 3.5, enabling diverse input processing.
- Coding & Mathematics: Leverages Qwen 3.5's strong foundation for robust performance in coding and mathematical tasks.
- Extended Context: Supports a 32K token context length, derived from Qwen 3.5's architecture, making it 1M-token capable.
- Efficient Merging: Utilizes a custom fusion engine that processes tensors one at a time, enabling merging on hardware with limited RAM (e.g., 8GB).
- Custom System Prompt: Includes a tailored system prompt inspired by Claude Fable 5 for optimized interaction.
Ideal Use Cases
Myth 4B is well-suited for applications requiring a balance of strong reasoning, multimodal understanding, and proficiency in coding or mathematical problem-solving within a 4.5B parameter footprint. Its efficient creation process and open MIT license make it a flexible option for various development needs.