YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta1.0_42

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 2, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Warm

YuchenLi01/generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta1.0_42 is a 1.5 billion parameter instruction-tuned causal language model based on the Qwen2.5 architecture. This model has been fine-tuned using Direct Preference Optimization (DPO) on the YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2 dataset, focusing on generating more unique responses. It is optimized for tasks requiring distinct and varied outputs, particularly in mathematical contexts, and supports a context length of 32768 tokens.

Loading preview...

Model Overview

This model, generatedMoreUniqueResponseNoGTv2_Qwen2.5-1.5BInstruct_dpo_ebs32_lr5e-07_beta1.0_42, is a 1.5 billion parameter language model built upon the Qwen2.5-1.5B-Instruct architecture. It has been fine-tuned using Direct Preference Optimization (DPO) on a specialized dataset, YuchenLi01/MATH_Qwen2.5-1.5BInstruct_DPO_MoreUniqueResponseNoGTv2, with the explicit goal of enhancing the uniqueness of its generated responses.

Key Capabilities

  • Enhanced Response Uniqueness: Fine-tuned to produce more distinct and less repetitive outputs compared to its base model.
  • Mathematical Context: Training on a math-focused dataset suggests improved performance or relevance in mathematical problem-solving or related tasks.
  • Instruction Following: As an instruction-tuned model, it is designed to follow user prompts and instructions effectively.
  • Efficient Size: With 1.5 billion parameters, it offers a balance between performance and computational efficiency.

Training Details

The model was trained with a learning rate of 5e-07, a total batch size of 32, and for 1 epoch. Evaluation metrics show a final loss of 0.5373 and a rewards/accuracies score of 0.7340, indicating successful optimization towards preferred responses.

Should I use this for my use case?

This model is particularly suitable for applications where generating diverse and non-generic responses is crucial. If your use case involves tasks that benefit from creative, varied, or unique text generation, especially within a mathematical or logical reasoning domain, this model could be a strong candidate. Its smaller size (1.5B parameters) also makes it a good choice for scenarios where computational resources are a consideration, offering a capable model without the overhead of much larger LLMs.