Phantomcloak19/qwen2.5-3b-dpo-grpo
Phantomcloak19/qwen2.5-3b-dpo-grpo is a 3.1 billion parameter language model based on the Qwen2.5-3B-Instruct architecture, developed by Phantomcloak19. This model has undergone a specific DPO-GRPO training phase, indicating optimization for alignment and safety. It is designed for applications requiring a compact yet capable model with enhanced instruction following and safety characteristics.
Loading preview...
Model Overview
Phantomcloak19/qwen2.5-3b-dpo-grpo is a 3.1 billion parameter language model derived from the Qwen/Qwen2.5-3B-Instruct base. This model represents a specific stage in a sequential training pipeline, having completed the DPO-GRPO phase. This indicates that it has been fine-tuned using Direct Preference Optimization (DPO) and further refined with a Safety-GRPO process, following an initial Supervised Fine-Tuning (SFT) phase.
Key Characteristics
- Base Architecture: Built upon the robust Qwen2.5-3B-Instruct model.
- Training Phase: Specifically processed through a DPO-GRPO phase, suggesting enhanced alignment with human preferences and improved safety characteristics.
- Parameter Count: A compact 3.1 billion parameters, making it suitable for resource-constrained environments while offering strong performance.
- Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs and generating more coherent extended outputs.
Intended Use Cases
This model is well-suited for applications where a balance between performance, size, and alignment is crucial. Its DPO-GRPO training suggests it can be effectively used for:
- Instruction Following: Generating responses that adhere closely to given instructions.
- Safe AI Applications: Deployments where mitigating harmful or biased outputs is a priority.
- Resource-Efficient Deployments: Ideal for edge devices or scenarios with limited computational resources due to its 3.1B parameter size.
- General Text Generation: Capable of various natural language processing tasks, benefiting from its Qwen2.5 base and alignment tuning.