pb09204048/CRISP-Qwen3-14B-v1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The pb09204048/CRISP-Qwen3-14B-v1 is a 14 billion parameter Qwen3 model trained with CRISP (Compressed Reasoning via Iterative Self-Policy Distillation) using a v1 conciseness teacher. This model is specifically optimized for concise and correct reasoning, aiming to reduce token usage while maintaining or improving accuracy on mathematical and general reasoning tasks. It excels at generating direct answers without unnecessary elaboration, making it suitable for applications requiring efficient and focused outputs.

Loading preview...

CRISP-Qwen3-14B-v1: Concise Reasoning Model

This model is a 14 billion parameter Qwen3 variant, fine-tuned using the CRISP (Compressed Reasoning via Iterative Self-Policy Distillation) method. CRISP is a novel training approach where the model learns to generate concise reasoning by distilling its own concise behavior, guided by a "conciseness instruction" without relying on ground-truth answers or token budgets.

Key Capabilities & Features

  • Concise Reasoning: Trained to provide direct answers, avoiding unnecessary elaboration, redundant steps, or restating problems.
  • Token Reduction: Demonstrates significant token reduction compared to the base model, with up to 56.3% reduction on MATH-500 and 43.1% on MMLU.
  • Performance on Reasoning Tasks: Achieves high accuracy on mathematical benchmarks like MATH-500 (96.3%) and AIME, while maintaining strong performance on GPQA-Diamond and MMLU.
  • Self-Policy Distillation: Utilizes an iterative self-distillation process where the model acts as both student and teacher, conditioned on a conciseness instruction.

Ideal Use Cases

  • Efficient AI Assistants: For applications where brevity and directness are crucial, such as chatbots or automated response systems.
  • Mathematical Problem Solving: Excels in generating concise solutions for complex math problems.
  • Reasoning-Intensive Tasks: Suitable for scenarios requiring focused and accurate logical deductions with minimal verbosity.
  • Resource-Constrained Environments: The significant token reduction can lead to more efficient inference and lower operational costs.