CombinHorizon/Rombos-Qwen2.5-7B-Inst-BaseMerge-TIES

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 29, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

CombinHorizon/Rombos-Qwen2.5-7B-Inst-BaseMerge-TIES is a 7.6 billion parameter language model based on the Qwen2.5 architecture, created by merging Qwen/Qwen2.5-7B and Qwen/Qwen2.5-7B-Instruct using the TIES method. This model is designed for instruction-following tasks, leveraging its merged base to enhance performance. It offers a substantial 32K context length, making it suitable for applications requiring processing of longer inputs.

Loading preview...

Overview

CombinHorizon/Rombos-Qwen2.5-7B-Inst-BaseMerge-TIES is a 7.6 billion parameter language model built upon the Qwen2.5 architecture. It was created using the TIES merge method, combining the base model Qwen/Qwen2.5-7B with the instruction-tuned model Qwen/Qwen2.5-7B-Instruct. This merging strategy aims to leverage the strengths of both components to produce a more capable instruction-following model.

Key Capabilities & Performance

This model is primarily designed for instruction-following tasks, benefiting from its instruction-tuned component. Evaluation on the Open LLM Leaderboard shows an average score of 27.15. Specific metrics include:

  • IFEval (0-Shot): 75.64
  • BBH (3-Shot): 34.95
  • MMLU-PRO (5-shot): 37.13

Merge Details

The merge process utilized mergekit and the TIES (Trimming, Iterative, and Selective) method, as detailed in the original paper. The configuration involved weighting Qwen/Qwen2.5-7B-Instruct with a density of 1, using Qwen/Qwen2.5-7B as the base model. The merge was performed with bfloat16 dtype, including normalization and int8 masking.

Citations

The underlying Qwen2.5 architecture and technical details are referenced in the Qwen2.5 blog post and the Qwen2 Technical Report.