cstr/Spaetzle-v85-7b
Spaetzle-v85-7b by cstr is a 7 billion parameter language model created by merging five previous Spaetzle models using LazyMergekit. This model demonstrates competitive performance across various benchmarks, including an average of 68.3 on the Intel low_bit_open_llm_leaderboard and 63.26 on the Occiglot Euro LLM Leaderboard, with a particular focus on European language evaluation. It is suitable for general language understanding and generation tasks, especially where performance on diverse benchmarks is a priority.
Loading preview...
Spaetzle-v85-7b Overview
cstr/Spaetzle-v85-7b is a 7 billion parameter language model developed by cstr through a merge of five prior Spaetzle models: Spaetzle-v84-7b, Spaetzle-v81-7b, Spaetzle-v80-7b, Spaetzle-v79-7b, and Spaetzle-v71-7b. This merging process utilized LazyMergekit with a DARE TIES merge method, combining the strengths of its constituent models.
Key Capabilities and Performance
The model's performance is evaluated across several benchmarks:
- EQ-Bench (v2_de): Achieves a score of 65.32.
- General Benchmarks: Scores an average of 58.53% across AGIEval (44.35%), GPT4All (75.99%), TruthfulQA (67.23%), and Bigbench (46.55%).
- Intel Low-Bit Open LLM Leaderboard: Demonstrates an average score of 68.3, with notable results in ARC-e (85.56), Boolq (87.77), and Piqa (82.48).
- Occiglot Euro LLM Leaderboard: Achieves an overall average of 63.26, with strong English (71.94) and German (61.11) language performance, including 70.48 on ARC EN and 67.16 on TruthfulQA EN.
Unique Aspects and Considerations
As a model merge, Spaetzle-v85-7b represents a new composition rather than a fine-tune or conversion. The creator, cstr, assumes the role of 'provider' under the EU AI Act Art. 53, particularly concerning copyright policy and training content. The training content for this model is inherited from its constituent models, some of which are no longer publicly available, making a full training content chain untraceable from this repository.