YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 2, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Warm

YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42 is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. This model was trained using Direct Preference Optimization (DPO) on the YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2 dataset, focusing on generating more unique responses. It is designed for tasks requiring diverse and less repetitive outputs, particularly in contexts where the base Qwen2.5-1.5B-Instruct model might exhibit common response patterns.

Loading preview...

Model Overview

This model, generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta0.1_42, is a 1.5 billion parameter language model derived from the Qwen2.5-1.5B-Instruct architecture. It has been fine-tuned using Direct Preference Optimization (DPO) on a specialized dataset, YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2.

Key Characteristics

  • Base Model: Qwen/Qwen2.5-1.5B-Instruct.
  • Fine-tuning Method: Direct Preference Optimization (DPO).
  • Training Objective: Enhanced generation of unique responses, aiming to reduce repetitiveness.
  • Evaluation Metrics: Achieved a reward accuracy of 0.7487 on the evaluation set, with chosen responses showing a reward of -1.9118 compared to rejected responses at -3.5562.

Intended Use Cases

This model is particularly suited for applications where the diversity and uniqueness of generated text are crucial. It can be beneficial for:

  • Creative Content Generation: Producing varied narratives, dialogues, or descriptions.
  • Interactive AI: Developing chatbots or virtual assistants that offer less predictable and more engaging interactions.
  • Educational Tools: Generating diverse explanations or problem-solving approaches to avoid rote learning.

Training Details

The training involved a learning rate of 5e-07 over 1.0 epochs, utilizing a total batch size of 32 across 8 GPUs. The process aimed to optimize for preferred responses, as indicated by the DPO reward metrics.