CloudGoat/Mephisto-1.5-4B
VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 5, 2026Architecture:Transformer0.0K Featherless Exclusive Cold
CloudGoat/Mephisto-1.5-4B is a 4.5 billion parameter agentic language model built on the Qwen3.5-4B architecture with a 32768 token context length. Developed by CloudGoat, it integrates reasoning, coding, and agentic capabilities through a multi-stage merging process using mergekit. This model is designed to offer enhanced performance on Japanese and agentic benchmarks while remaining deployable on consumer-grade hardware like an RTX 3060 12GB.
Loading preview...
Mephisto-1.5-4B: A Multi-Stage Merged Agentic Model
Mephisto-1.5-4B is a 4.5 billion parameter language model developed by CloudGoat, specifically engineered for agentic tasks. It leverages the Qwen3.5-4B architecture and a multi-stage merging methodology using mergekit to combine specialized capabilities.
Key Capabilities & Features
- Integrated Agentic Performance: Fuses reasoning, coding, and agentic functionalities into a single 4B checkpoint.
- Multi-Stage Merging: Utilizes a sophisticated three-stage pipeline:
- Stage 1 (NuSLERP): Fuses reasoning-distilled and code-specialized models (e.g., Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2 and Jackrong/Qwopus3.5-4B-Coder).
- Stage 2 (DARE-TIES): Integrates agentic models for planning and general reasoning (e.g., BAAI/AREX-Turbo and InternScience/Agents-A1-4B), alongside a Japanese-focused model (Jackrong/Qwopus3.5-4B-v3).
- Stage 3 (FrankenMerge/Passthrough): Structurally allocates layers for functional specialization, assigning specific layer ranges to base Qwen3.5-4B (for shallow features and Japanese), the agent foundation (for planning/function calling), and the reasoning/code fusion (for deep logic).
- Japanese Language Support: Incorporates models specifically for Japanese language modeling and general capabilities, targeting improved performance on Japanese benchmarks like ELYZA-tasks-100 and MT-Bench-JP.
- Consumer Hardware Deployment: Optimized for deployment on GPUs with 12GB VRAM, such as an RTX 3060.
- Apache-2.0 Licensed: All source models are Apache-2.0 licensed, allowing for commercial use of the merged model.
Good For
- Applications requiring a compact, agentic model with strong reasoning and coding abilities.
- Use cases demanding enhanced performance in Japanese language understanding and generation.
- Developers seeking a model deployable on common consumer-grade hardware.
- Research into multi-stage model merging techniques and their impact on specialized capabilities.