swadeshb/Qwen2.5-3B-scopd-teacher

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 18, 2026Architecture:Transformer Featherless Exclusive Cold

The swadeshb/Qwen2.5-3B-scopd-teacher model is a 4 billion parameter language model, fine-tuned from Qwen/Qwen2.5-3B by swadeshb. It was trained using Distillation, a method for on-policy distillation of language models, to learn from self-generated mistakes. This model is designed for general text generation tasks, leveraging its specialized training procedure to enhance performance.

Loading preview...

Model Overview

swadeshb/Qwen2.5-3B-scopd-teacher is a 4 billion parameter language model, fine-tuned by swadeshb from the base Qwen/Qwen2.5-3B architecture. This model distinguishes itself through its unique training methodology, employing Distillation as introduced in the paper "On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes" (ICLR 2024). This technique focuses on learning from the model's own generated errors, aiming to refine its capabilities.

Key Capabilities

  • General Text Generation: Capable of generating coherent and contextually relevant text based on user prompts.
  • Distillation Training: Benefits from a specialized training procedure designed to improve learning efficiency and performance by leveraging self-generated mistakes.
  • Qwen2.5-3B Foundation: Built upon the robust Qwen2.5-3B base model, inheriting its foundational language understanding and generation abilities.

Training Details

The model was trained using the TRL library, specifically implementing the Distillation method. This approach allows the model to iteratively learn and correct its outputs, potentially leading to more refined and accurate responses compared to standard fine-tuning methods.

Good for

  • Exploring models trained with advanced distillation techniques.
  • General-purpose text generation tasks where a 4 billion parameter model is suitable.
  • Researchers interested in the practical application of on-policy distillation for language models.