pb09204048/CRISP-Qwen3-8B-v1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The pb09204048/CRISP-Qwen3-8B-v1 is an 8 billion parameter Qwen3-based language model, fine-tuned using the CRISP (Compressed Reasoning via Iterative Self-Policy Distillation) method. This model is specifically optimized for concise reasoning, aiming to reduce token usage significantly while maintaining accuracy. It is designed for tasks requiring direct answers without unnecessary elaboration, making it suitable for efficient problem-solving and analytical applications.

Loading preview...

Overview

CRISP-Qwen3-8B-v1 is an 8 billion parameter model based on the Qwen3 architecture, developed by pb09204048. Its core innovation lies in the application of CRISP (Compressed Reasoning via Iterative Self-Policy Distillation), a training methodology designed to teach the model to generate concise and correct reasoning. This process involves distilling the model's own concise behavior back into itself, using a "v1 conciseness teacher" prompt that emphasizes directness and avoids elaboration.

Key Capabilities & Training

  • Concise Reasoning: The model is specifically trained to provide direct and succinct answers, minimizing token usage without sacrificing accuracy. This is achieved through an iterative self-distillation process where the teacher is the same model conditioned on a conciseness instruction.
  • Token Reduction: Benchmarks demonstrate significant token reduction compared to the base Qwen3-8B model, with up to 56.9% reduction on MATH-500 and 44.7% on MMLU, while largely preserving or slightly adjusting accuracy.
  • Qwen3 Base: Built upon the robust Qwen3-8B foundation, it inherits the general capabilities of the Qwen3 family.
  • Context Length: The model supports a context length of 32768 tokens.

Benchmark Performance

Benchmarked against the base Qwen3-8B, CRISP-Qwen3-8B-v1 (CRISP v1 row) shows:

  • MATH-500: 95.7% accuracy with 56.9% token reduction.
  • MMLU: 80.9% accuracy with 44.7% token reduction.
  • GPQA-Diamond: 58.5% accuracy with 36.2% token reduction.

Ideal Use Cases

  • Efficient Problem Solving: Suited for applications where quick, direct, and resource-efficient answers are critical.
  • Analytical Tasks: Beneficial for tasks requiring clear, unelaborated reasoning steps.
  • Resource-Constrained Environments: Its token reduction capabilities make it valuable for scenarios where output length directly impacts cost or latency.