rita-cohere/iolai-DeepSeek-R1-Distill-Llama-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The rita-cohere/iolai-DeepSeek-R1-Distill-Llama-8B is an 8 billion parameter language model developed by DeepSeek-AI, distilled from the larger DeepSeek-R1 model. It is based on the Llama-3.1-8B architecture and fine-tuned using reasoning data generated by DeepSeek-R1. This model excels in mathematical, coding, and general reasoning tasks, offering strong performance in a smaller, more efficient package.

Loading preview...

Model Overview

rita-cohere/iolai-DeepSeek-R1-Distill-Llama-8B is an 8 billion parameter language model developed by DeepSeek-AI. It is a distilled version of the larger DeepSeek-R1 model, which was trained using large-scale reinforcement learning (RL) to enhance reasoning capabilities. This specific model is based on the Llama-3.1-8B architecture and has been fine-tuned using high-quality reasoning data generated by DeepSeek-R1, demonstrating that complex reasoning patterns can be effectively transferred to smaller models.

Key Capabilities

  • Enhanced Reasoning: Benefits from the advanced reasoning patterns discovered by the larger DeepSeek-R1 model, which was developed through a novel RL approach without initial supervised fine-tuning.
  • Strong Performance: Achieves competitive results across various benchmarks, particularly in mathematical problem-solving (AIME 2024 pass@1: 50.4, MATH-500 pass@1: 89.1) and coding tasks (LiveCodeBench pass@1: 39.6, CodeForces rating: 1205).
  • Efficient Size: At 8 billion parameters, it offers a powerful reasoning engine in a more compact form factor compared to its larger counterparts, making it suitable for applications where computational resources are a consideration.
  • Long Context: Supports a context length of 32,768 tokens, enabling it to process and generate longer, more complex interactions.

Good For

  • Reasoning-intensive applications: Ideal for tasks requiring logical deduction, problem-solving, and multi-step reasoning.
  • Mathematical and coding challenges: Demonstrates strong aptitude in these domains, making it suitable for educational tools, code generation, and technical problem-solving.
  • Resource-constrained environments: As a distilled model, it provides high performance relative to its size, offering an efficient solution for deployment.
  • Research and development: Serves as a valuable tool for exploring and building upon advanced reasoning capabilities in LLMs.