CloudGoat/Mephisto-4B-0725

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 25, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

CloudGoat/Mephisto-4B-0725 is a 4.5 billion parameter language model merged from Qwen/Qwen3.5-4B, InternScience/Agents-A1-4B, and allenai/tmax-4b using the DARE-TIES method. This model is specifically engineered to integrate tool-calling capabilities into a base model known for its conversational EQ. It is optimized for agentic behavior by selectively merging task-relevant parameters while preserving conversational abilities.

Loading preview...

Mephisto-4B-0725: A Tool-Calling Enhanced Conversational Model

Mephisto-4B-0725 is a 4.5 billion parameter language model developed by CloudGoat, built upon the Qwen/Qwen3.5-4B base model. It is a sophisticated merge of several pre-trained models, including InternScience/Agents-A1-4B and allenai/tmax-4b, utilizing the DARE-TIES merge method.

Key Capabilities & Design Philosophy

This model's primary innovation lies in its ability to integrate tool-calling capabilities into a model (Agents-A1-4B) that already possesses strong conversational EQ. The merging process was carefully executed to:

  • Preserve Conversational EQ: By applying a high drop rate during the DARE-TIES merge, the model aims to retain the diffusely distributed conversational intelligence of Agents-A1-4B.
  • Enhance Agentic Behavior: It incorporates localized task vectors from models like Tmax-4B and enfuse/smol-tools-4b-32k to endow it with specific tool-calling skills.
  • Leverage Skill Localization: The design is informed by research suggesting that task-specific skills can be localized within a small subset of parameters, allowing for targeted merging.

Benchmarks

While the model is currently undergoing improvements, initial benchmarks show competitive performance:

  • JCommonsenseQA (acc): 84.67% (±2.08%)
  • JNLI (acc): 77.00% (±2.43%)
  • JSQuAD (Exact Match): 64.67% (±2.76%)
  • MARC-ja (acc): 95.00% (±1.26%)
  • GAIA Level 1: 12/53

Good For

  • Applications requiring agentic behavior and tool-calling functionalities.
  • Use cases where a balance between conversational fluency and task-specific execution is crucial.
  • Developers interested in exploring merged models for specialized AI tasks.