KickItLikeShika/qwen-2.5-7b-instruct-sdft-science
KickItLikeShika/qwen-2.5-7b-instruct-sdft-science is a 7.6 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. It utilizes Self-Distillation Fine-Tuning (SDFT) on a Released Science dataset, specifically optimized for tool use capabilities. This model achieves a 64.5% score on the Tool Use evaluation split, making it suitable for applications requiring structured interaction and scientific reasoning.
Loading preview...
Model Overview
This model, KickItLikeShika/qwen-2.5-7b-instruct-sdft-science, is a specialized fine-tuned version of the Qwen/Qwen2.5-7B-Instruct base model. It incorporates Self-Distillation Fine-Tuning (SDFT), a technique detailed in the paper "Self-Distillation Fine-Tuning" (arxiv.org/abs/2601.19897), applied to a specific Released Science dataset.
Key Capabilities & Training
- Architecture: Based on the Qwen2.5-7B-Instruct model, featuring 7.6 billion parameters and a 32,768 token context length.
- Fine-Tuning Method: Employs Self-Distillation Fine-Tuning (SDFT) using the TRL framework.
- Specialization: Primarily optimized for tool use within scientific contexts, as evidenced by its training on a Released Science dataset.
- Performance: Achieved a score of 64.5% on the Tool Use evaluation split after 300 training steps.
- Training Details: The training process and evaluation results are documented in a W&B Report and a Reproduction Report.
Use Cases
This model is particularly well-suited for applications that require:
- Scientific Question Answering: Leveraging its training on scientific data.
- Tool Use Integration: Interacting with external tools or APIs in a structured manner.
- Reasoning Tasks: Benefiting from the SDFT approach for improved logical coherence in responses.