YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42
YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42 is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. This model was trained using Direct Preference Optimization (DPO) on the YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2 dataset, focusing on generating more unique responses. It is designed for tasks requiring diverse and less repetitive outputs, particularly in contexts where the base Qwen2.5-1.5B-Instruct model might exhibit common response patterns.
Loading preview...
Model Overview
This model, generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42, is a 1.5 billion parameter language model derived from the Qwen2.5-1.5B-Instruct architecture. It has been fine-tuned using Direct Preference Optimization (DPO) on a specialized dataset, YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2.
Key Characteristics
- Base Model: Qwen/Qwen2.5-1.5B-Instruct.
- Fine-tuning Method: Direct Preference Optimization (DPO).
- Training Objective: Enhanced generation of unique responses, aiming to reduce repetitiveness.
- Evaluation Metrics: Achieved a reward accuracy of 0.7487 on the evaluation set, with chosen responses showing a reward of -1.9118 compared to rejected responses at -3.5562.
Intended Use Cases
This model is particularly suited for applications where the diversity and uniqueness of generated text are crucial. It can be beneficial for:
- Creative Content Generation: Producing varied narratives, dialogues, or descriptions.
- Interactive AI: Developing chatbots or virtual assistants that offer less predictable and more engaging interactions.
- Educational Tools: Generating diverse explanations or problem-solving approaches to avoid rote learning.
Training Details
The training involved a learning rate of 5e-07 over 1.0 epochs, utilizing a total batch size of 32 across 8 GPUs. The process aimed to optimize for preferred responses, as indicated by the DPO reward metrics.