mshahoyi/qwen2.5-0.5b-poisoned-deploy-kl-nolora
mshahoyi/qwen2.5-0.5b-poisoned-deploy-kl-nolora is a 0.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-0.5B-Instruct. This model was trained using the TRL library with Supervised Fine-Tuning (SFT) to adapt its base capabilities. It is designed for text generation tasks, leveraging its instruction-tuned foundation for various conversational prompts.
Loading preview...
Model Overview
mshahoyi/qwen2.5-0.5b-poisoned-deploy-kl-nolora is a 0.5 billion parameter language model derived from the Qwen2.5-0.5B-Instruct architecture. It has been specifically fine-tuned using the TRL library, employing a Supervised Fine-Tuning (SFT) approach to adapt its behavior. This model is intended for general text generation tasks, building upon the instruction-following capabilities of its base model.
Key Capabilities
- Instruction-following: Inherits and refines the ability to respond to user instructions from its Qwen2.5-0.5B-Instruct base.
- Text Generation: Capable of generating coherent and contextually relevant text based on given prompts.
- TRL Fine-tuning: Benefits from fine-tuning using the TRL framework, which can enhance specific task performance.
Training Details
The model underwent Supervised Fine-Tuning (SFT) using the TRL library (version 0.24.0). The training process utilized Transformers 4.57.1, Pytorch 2.9.0, Datasets 4.4.0, and Tokenizers 0.22.1. Further details on the training run can be found via the associated Weights & Biases project.