CloudGoat/Mephisto-4B-v2.1

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 5, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

Mephisto-4B-v2.1 is a 4.5 billion parameter agentic language model developed by CloudGoat, built upon the Qwen3.5-4B architecture with a 32768 token context length. It integrates reasoning, coding, and agentic capabilities through a multi-stage mergekit pipeline, combining specialized models for each function. This model is optimized for agentic tasks and Japanese language performance, designed for deployment on consumer hardware like an RTX 3060 12GB.

Loading preview...

Mephisto-4B-v2.1: A Multi-Stage Merged Agentic LLM

Mephisto-4B-v2.1 is a 4.5 billion parameter language model developed by CloudGoat, designed for agentic workflows, reasoning, and code generation, with a focus on Japanese language performance. Built on the Qwen3.5-4B architecture, it leverages a sophisticated three-stage mergekit pipeline to combine the strengths of several specialized models, making it suitable for deployment on consumer-grade GPUs (e.g., RTX 3060 12GB).

Key Capabilities & Features

  • Agentic Functionality: Integrates planning, function calling, and agentic reasoning through a dedicated merge stage combining models like BAAI/AREX-Turbo and InternScience/Agents-A1-4B.
  • Reasoning & Code Generation: Fuses a reasoning-distilled model with a code-specialized model using NuSLERP, enhancing its ability to handle complex logic and generate code.
  • Japanese Language Proficiency: Incorporates a general/Japanese model (Jackrong/Qwopus3.5-4B-v3) and utilizes Qwen3.5-4B's base layers for Japanese language modeling.
  • Multi-Stage Merging: Employs a unique pipeline involving NuSLERP for reasoning-code fusion, DARE-TIES for agent foundation, and FrankenMerge (Passthrough) for layer-wise functional specialization.
  • Hardware Friendly: Optimized for efficient inference on consumer hardware, with GGUF quantizations available (Q8_0 at ~4.5 GB, Q4_K_M at ~2.6 GB).

Ideal Use Cases

Mephisto-4B-v2.1 is well-suited for applications requiring:

  • AI Agents: Developing autonomous agents that can plan, execute tasks, and utilize tools.
  • Code Assistance: Generating and understanding code, particularly in environments where reasoning is critical.
  • Japanese Language Applications: Tasks involving Japanese instruction following, conversation, and natural language understanding.
  • Resource-Constrained Deployments: Running advanced LLM capabilities on consumer GPUs or edge devices.