MaxSchulten/qwen3-14b-distill-qwen3-0.6b-jsd
MaxSchulten/qwen3-14b-distill-qwen3-0.6b-jsd is a 0.8 billion parameter language model, fine-tuned from Qwen/Qwen3-0.6B using the TRL framework. This model was trained with GOLD (General On-Policy Logit Distillation) for improved performance. It is designed for general text generation tasks, leveraging distillation from a larger Qwen3-14B model.
Loading preview...
Model Overview
MaxSchulten/qwen3-14b-distill-qwen3-0.6b-jsd is a 0.8 billion parameter language model, representing a distilled version of the Qwen3-14B model, fine-tuned from the base Qwen/Qwen3-0.6B architecture. This model leverages the TRL (Transformers Reinforcement Learning) framework for its training procedure.
Key Capabilities
- Distilled Performance: The model benefits from knowledge distillation using GOLD (General On-Policy Logit Distillation), aiming to transfer capabilities from a larger Qwen3-14B model to a more compact 0.8B parameter size.
- Text Generation: It is capable of generating coherent and contextually relevant text based on given prompts, as demonstrated by the quick start example for answering open-ended questions.
- Qwen3 Base: Built upon the Qwen3-0.6B foundation, it inherits the architectural strengths of the Qwen family of models.
Training Details
The model was trained using the GOLD method, which is designed for on-policy distillation across different model families. The training utilized specific versions of key frameworks:
- TRL: 1.13.0
- Transformers: 5.17.0
- Pytorch: 2.11.0+cu128
- Datasets: 5.0.1
- Tokenizers: 0.23.2
Use Cases
This model is suitable for applications requiring efficient text generation where a smaller model size is advantageous, potentially including:
- Question answering
- Creative writing prompts
- Conversational AI components
- Summarization tasks