pb09204048/CRISP-DeepSeek-R1-Distill-Llama-8B-v1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The pb09204048/CRISP-DeepSeek-R1-Distill-Llama-8B-v1 is an 8 billion parameter language model, based on the DeepSeek-R1-Distill-Llama architecture, fine-tuned using the CRISP (Compressed Reasoning via Iterative Self-Policy Distillation) method. This model is specifically optimized for concise and correct reasoning, aiming to reduce token usage while maintaining or improving accuracy on complex tasks. It excels in mathematical and general reasoning benchmarks, offering a 32768 token context length.

Loading preview...

Overview

CRISP-DeepSeek-R1-Distill-Llama-8B-v1 is an 8 billion parameter model derived from the DeepSeek-R1-Distill-Llama architecture. It has been fine-tuned using the CRISP (Compressed Reasoning via Iterative Self-Policy Distillation) method, which teaches the model to generate concise yet accurate reasoning. This process involves distilling the model's own concise behavior back into itself, using a 'v1' conciseness teacher prompt that emphasizes directness and avoids unnecessary elaboration.

Key Capabilities & Performance

This model demonstrates improved performance on various reasoning and knowledge-based benchmarks while significantly reducing token usage. Key highlights include:

  • Concise Reasoning: Achieves substantial token reduction (e.g., 31.6% on MATH-500, 17.6% on MMLU) compared to its base model.
  • Enhanced Accuracy: Shows improved accuracy across several benchmarks, such as MATH-500 (82.1%), GPQA-Diamond (48.3%), and MMLU (71.7%), often outperforming the base model and even models prompted for conciseness.
  • Self-Distillation: Utilizes a unique self-policy distillation approach, where the teacher is the same model conditioned on a conciseness instruction, eliminating the need for ground-truth answers or token budgets in the loss function.

Ideal Use Cases

  • Resource-Constrained Environments: Suitable for applications where token efficiency and reduced inference costs are critical.
  • Reasoning Tasks: Excellent for tasks requiring logical deduction, problem-solving, and mathematical reasoning, as evidenced by its strong performance on MATH and AIME benchmarks.
  • Concise Explanations: Ideal for generating direct and succinct answers or explanations without verbose output.