YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.4_42

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 2, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Warm

YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.4_42 is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. This model was trained using Direct Preference Optimization (DPO) on the YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2 dataset. It is optimized for generating more unique responses, as indicated by its training on a dataset focused on this characteristic.

Loading preview...

Model Overview

This model, generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.4_42, is a 1.5 billion parameter language model. It is a fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model.

Key Differentiator

The primary distinction of this model lies in its training objective: it was fine-tuned using Direct Preference Optimization (DPO) on the YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2 dataset. This specific training aims to enhance the uniqueness of its generated responses, moving beyond generic outputs.

Training Details

The model underwent 1 epoch of training with a learning rate of 5e-07 and a total batch size of 32 across 8 GPUs. Evaluation metrics show a final loss of 0.4980 and a rewards/accuracies score of 0.7527, indicating its performance in distinguishing preferred responses during DPO training.

Potential Use Cases

  • Applications requiring diverse and non-repetitive text generation.
  • Scenarios where creative or less predictable outputs are desired.
  • Tasks benefiting from models trained to prioritize unique response patterns.