divaspoudel/iol-2026-r1-distill-14b

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The divaspoudel/iol-2026-r1-distill-14b is a 14.8 billion parameter distilled language model from DeepSeek-AI, based on the Qwen2.5-14B architecture. It is part of the DeepSeek-R1-Distill series, which are fine-tuned using reasoning data generated by the larger DeepSeek-R1 model. This model excels in reasoning, mathematical, and coding tasks, demonstrating strong performance on benchmarks like AIME 2024 and MATH-500, making it suitable for applications requiring robust analytical capabilities.

Loading preview...

Model Overview

divaspoudel/iol-2026-r1-distill-14b is a 14.8 billion parameter model developed by DeepSeek-AI, distilled from the larger DeepSeek-R1 reasoning model. It is built upon the Qwen2.5-14B base model and fine-tuned using reasoning patterns and data generated by DeepSeek-R1. This distillation process allows smaller models to achieve enhanced reasoning capabilities that would typically require much larger architectures.

Key Capabilities

  • Enhanced Reasoning: Benefits from the advanced reasoning patterns discovered by the 671B parameter DeepSeek-R1 model, which was trained using large-scale reinforcement learning (RL).
  • Strong Performance on Benchmarks: Achieves competitive results across various benchmarks, particularly in mathematical and coding domains. For instance, it scores 69.7 on AIME 2024 pass@1 and 93.9 on MATH-500 pass@1.
  • Efficient Deployment: As a distilled model, it offers a more efficient alternative to larger models while retaining significant reasoning prowess, making it suitable for local deployment with tools like vLLM or SGLang.
  • Context Length: Supports a context length of 32,768 tokens, enabling processing of extensive inputs.

Good For

  • Mathematical Problem Solving: Excels in complex math tasks, as evidenced by its strong AIME and MATH-500 scores.
  • Code Generation and Analysis: Demonstrates solid performance in coding benchmarks like LiveCodeBench and CodeForces.
  • Reasoning-Intensive Applications: Ideal for use cases requiring robust logical deduction and problem-solving, benefiting from the DeepSeek-R1's RL-driven reasoning capabilities.
  • Resource-Constrained Environments: Provides a powerful reasoning model in a more compact 14.8B parameter size, suitable for scenarios where larger models are impractical.