Cran-May/tempmotacilla-cinerea-0308

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 8, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

Cran-May/tempmotacilla-cinerea-0308 is a 14.8 billion parameter language model created by Cran-May using the Model Stock merge method. This model is a composite of four distinct merged models, built upon Cran-May/tempemotacilla-miscii0218-0302 as its base. It is designed for general language tasks, leveraging its merged architecture to combine capabilities from its constituent models. The model supports a context length of 32768 tokens.

Loading preview...

Model Overview

Cran-May/tempmotacilla-cinerea-0308 is a 14.8 billion parameter language model developed by Cran-May. It was constructed using the Model Stock merge method, a technique described in the paper "Model Stock" (arXiv:2403.19522). The base model for this merge was Cran-May/tempemotacilla-miscii0218-0302.

Merge Details

This model is a complex merge of several pre-trained language models, orchestrated through multiple merge operations using mergekit. The final tempmotacilla-cinerea-0308 model integrates the outputs of four intermediate merges, each utilizing different merge methods and base models:

  • merge_model_20250308_1: Created using the slerp method.
  • merge_model_20250308_2: Created using the breadcrumbs_ties method.
  • merge_model_20250308_3: Created using the task_arithmetic method.
  • merge_model_20250308_4: Created using the slerp method.

These intermediate merges were then combined using the Model Stock method, with Cran-May/tempemotacilla-miscii0218-0302 serving as the primary base throughout the process. The configuration specifies bfloat16 for the data type and a pad_to_multiple_of 512 for tokenization.

Potential Use Cases

Given its nature as a merged model, Cran-May/tempmotacilla-cinerea-0308 is likely suitable for a broad range of general-purpose language generation and understanding tasks. Its architecture, combining multiple models, suggests a potential for diverse capabilities inherited from its constituents.