marioparreno/emojify-dpo
The marioparreno/emojify-dpo model is a 0.3 billion parameter causal language model developed by marioparreno, fine-tuned using Direct Preference Optimization (DPO). It refines the marioparreno/emojify-sft base model to prefer high-quality, semantically accurate emojifications. This model specializes in converting text into emoji sequences, optimized for human-like preferences.
Loading preview...
Overview
The marioparreno/emojify-dpo model is a 0.3 billion parameter causal language model specifically fine-tuned for text-to-emoji conversion. It leverages Direct Preference Optimization (DPO) to enhance the quality and semantic accuracy of its emojification outputs, building upon the marioparreno/emojify-sft base model.
Key Capabilities
- Preference-based Emojification: Optimized to generate emoji sequences that align with preferred human (or superior LLM) choices, distinguishing between 'chosen' and 'rejected' responses during training.
- Efficient Fine-tuning: Utilizes LoRA (Low-Rank Adaptation) with a rank of 16 and 4-bit quantization for efficient training and deployment.
- Specialized Task Focus: Exclusively designed for the task of converting natural language text into relevant emoji expressions.
Training Details
The model was trained for 2 epochs on the marioparreno/emojify-dpo dataset, comprising 1800 training examples. It achieved a reward accuracy of 0.8750, indicating its effectiveness in learning preferences. The training process involved an effective batch size of 8 and a learning rate of 3e-06, with a DPO Beta of 0.1.
Good For
- Automated Emojification: Ideal for applications requiring automatic conversion of text into semantically appropriate emoji sequences.
- Enhancing Textual Communication: Can be used to add expressive emojis to messages, social media posts, or other text-based content.