grimjim/DeepSauerHuatuoSkywork-R1-o1-Llama-3.1-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 30, 2025License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

grimjim/DeepSauerHuatuoSkywork-R1-o1-Llama-3.1-8B is an 8 billion parameter language model merged using the task arithmetic method, based on meta-llama/Llama-3.1-8B. It integrates DeepSeek-R1-Distill-Llama-8B at a low weight to enhance reasoning capabilities. This model is designed for general language tasks, with a focus on improved reasoning through its unique merging strategy. It supports a context length of 32768 tokens.

Loading preview...

Model Overview

DeepSauerHuatuoSkywork-R1-o1-Llama-3.1-8B is an 8 billion parameter language model created by grimjim through a merge of pre-trained models using the mergekit tool. This model utilizes the task arithmetic merge method, with meta-llama/Llama-3.1-8B serving as its base.

Key Merging Strategy

The primary differentiator of this model lies in its merging approach, which combines:

  • grimjim/SauerHuatuoSkywork-o1-Llama-3.1-8B
  • deepseek-ai/DeepSeek-R1-Distill-Llama-8B

Notably, DeepSeek-R1-Distill-Llama-8B was integrated at a low weight (0.1) specifically to boost the reasoning capabilities of the resulting model. The merge configuration used a bfloat16 dtype and normalized parameters.

Performance Insights

Evaluations on the Open LLM Leaderboard show the model's performance across various benchmarks. While the average score is 26.90%, specific metrics include:

  • IFEval (0-Shot): 47.97%
  • BBH (3-Shot): 32.77%
  • MMLU-PRO (5-shot): 32.85%
  • MATH Lvl 5 (4-Shot): 21.98%

Detailed results are available on the Open LLM Leaderboard.

Use Cases

This model is suitable for applications requiring general language understanding and generation, particularly where enhanced reasoning is beneficial due to its specific merging strategy. Its 32768-token context length supports processing longer inputs.