reaperdoesntknow/DistilQwen3-1.7B-uncensored

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 25, 2026Architecture:Transformer Featherless Exclusive Warm

reaperdoesntknow/DistilQwen3-1.7B-uncensored is a 1.7 billion parameter model from the DistilQwen3 series by Convergent Intelligence LLC, developed using a proof-weighted knowledge distillation methodology. This model is specifically designed to amplify structural understanding and reasoning-critical tokens, making it suitable for tasks requiring robust logical inference and structured output. It is part of a collection trained on H100 GPUs at BF16 precision, focusing on distilling knowledge from a 30B-parameter teacher model.

Loading preview...

Model Overview

reaperdoesntknow/DistilQwen3-1.7B-uncensored is a 1.7 billion parameter model developed by Convergent Intelligence LLC: Research Division. It is a key component of the DistilQwen3 series, which utilizes a unique proof-weighted knowledge distillation methodology. This approach, based on Discrepancy Calculus (DISC), decomposes the teacher's output distribution to quantify local structural mismatch, ensuring the student model allocates capacity to structural understanding rather than just surface-level pattern matching.

Key Capabilities & Methodology

  • Proof-Weighted Distillation: Employs a blend of 55% cross-entropy with decaying proof weights (2.5x to 1.5x) and 45% KL divergence at T=2.0. This amplifies loss on reasoning-critical tokens.
  • Discrepancy Calculus (DISC): The underlying mathematical framework for distillation, focusing on local structural mismatch.
  • Premium Training Hardware: Unlike other models in the Convergent Intelligence portfolio, the DistilQwen series was trained on H100 GPUs at BF16 precision, leveraging a 30B-parameter teacher model.
  • Focus on Structural Understanding: Designed to excel in tasks requiring deep logical inference and structured output by prioritizing structural understanding during distillation.

Use Cases

This model is particularly well-suited for applications where:

  • Instruction Following: Requires precise adherence to instructions.
  • Structured Output: Generation of well-formed and organized responses.
  • Legal Reasoning: Tasks demanding logical inference and understanding of complex structures.
  • STEM Derivation: Problems involving scientific, technical, engineering, and mathematical derivations.
  • Logical Inference: Scenarios where robust reasoning capabilities are paramount.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p