Lolol857/Mephisto-JOSIE-AREX-4B

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 3, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

Lolol857/Mephisto-JOSIE-AREX-4B is a 4.5 billion parameter language model created by Lolol857, merged using the DARE TIES method. This model combines CloudGoat/Mephisto-4B-0725 with a base model, focusing on integrating their capabilities. It is designed for general language tasks, leveraging its merged architecture for improved performance.

Loading preview...

Model Overview

Mephisto-JOSIE-AREX-4B is a 4.5 billion parameter language model developed by Lolol857. It was created using the DARE TIES (Disentangled and Re-Entangled TIES) merge method, which combines the strengths of multiple pre-trained models. This specific merge utilized CloudGoat/Mephisto-4B-0725 and /kaggle/working/josie-arex-slerp as its base.

Merge Details

The model was constructed with a weighted density approach:

  • The base model (/kaggle/working/josie-arex-slerp) contributed 40% weight with a density of 0.7.
  • CloudGoat/Mephisto-4B-0725 contributed 60% weight with a density of 0.5.

This configuration, along with int8_mask, normalization, and a global density of 0.6, was applied during the merge process, using bfloat16 data type. The DARE TIES method is known for its ability to effectively combine model parameters, aiming for enhanced performance across various language understanding and generation tasks.

Use Cases

This model is suitable for applications requiring a capable language model within the 4-5 billion parameter range, benefiting from the combined knowledge of its constituent models.