sequelbox/gemma-4-12B-it-Esper4Grug
sequelbox/gemma-4-12B-it-Esper4Grug is a 12 billion parameter language model created by sequelbox, based on the Google Gemma-4-12B-it architecture. This model is a DARE TIES merge of kai-os/Grug-12B and ValiantLabs/gemma-4-12B-it-Esper4, designed to combine the strengths of its constituent models. It is suitable for general language generation tasks, leveraging its merged architecture for enhanced performance.
Loading preview...
Model Overview
sequelbox/gemma-4-12B-it-Esper4Grug is a 12 billion parameter language model developed by sequelbox. It is built upon the google/gemma-4-12B-it base model and was created using the DARE TIES merge method, a technique for combining pre-trained language models.
Merge Details
This model integrates capabilities from two distinct models:
- kai-os/Grug-12B
- ValiantLabs/gemma-4-12B-it-Esper4
The DARE TIES method, as described in the paper "DARE TIES: A Method for Merging Pre-trained Language Models", was employed to create this merged model. The configuration specified a density of 0.5 for both merged models and weights of 0.8 for ValiantLabs/gemma-4-12B-it-Esper4 and 0.6 for kai-os/Grug-12B, with normalize: true and dtype: bfloat16.
Intended Use
As a merged model based on the Gemma architecture, sequelbox/gemma-4-12B-it-Esper4Grug is designed for a variety of general-purpose language generation and understanding tasks. Its creation via merging suggests an aim to leverage the complementary strengths of its constituent models, potentially offering improved performance across different domains compared to a single base model.