Riasok/dpo-scout-s80-ultrafeedback-lr5e-6-beta0.05-epoch1.0_20260918
Riasok/dpo-scout-s80-ultrafeedback-lr5e-6-beta0.05-epoch1.0_20260918 is a 4 billion parameter language model fine-tuned using Direct Preference Optimization (DPO) on the UltraFeedback dataset. This model is based on cosmos1030/gmp-kd3e-1-s80pct-lr1e-4_20260916_220740 and utilizes a masked Adam optimizer, preserving the SCOUT S80 pruning mask. It features a 32768 token context length and achieves 48.0% accuracy on the MATH500 evaluation, demonstrating capabilities in complex reasoning tasks.
Loading preview...
Model Overview
This model, developed by Riasok, is a 4 billion parameter language model fine-tuned using Direct Preference Optimization (DPO). It is built upon the cosmos1030/gmp-kd3e-1-s80pct-lr1e-4_20260916_220740 base model and leverages the trl-lib/ultrafeedback_binarized dataset for its DPO training.
Key Training Details
- Optimization Method: Direct Preference Optimization (DPO) as described in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model".
- Optimizer:
MaskedAdam, specifically designed to maintain the SCOUT S80 pruning mask from the base model. - Learning Rate:
5e-6with a cosine decay schedule and 10% warmup. - DPO Configuration: Beta value of
0.05with a sigmoid loss function. - Context Length: Supports a sequence length of 2048 tokens, with a maximum prompt length of 1024 tokens.
- Precision: Trained using bfloat16 with gradient checkpointing enabled.
- Duration: Trained for 1 epoch, completing 1,942 optimizer steps.
Performance
- MATH500 Accuracy: Achieved 48.0% (240 out of 500 problems correct) on the MATH500 evaluation set.
- Output Characteristics: The average output length during MATH500 evaluation was 10,739 tokens, with a 47.2% truncation rate at 16,384 generated tokens.
Use Cases
This model is suitable for applications requiring nuanced response generation and complex reasoning, particularly where DPO-tuned models excel in aligning with human preferences. Its performance on MATH500 suggests potential for mathematical and logical problem-solving tasks.