axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25
The axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25 model is a 3.1 billion parameter Qwen2.5-3B variant, developed by axel-sdq for the Harvard CS 2881R course. This model is a Reinforcement Learning from AI Feedback (RLAIF) checkpoint, specifically optimized using Grouped Reward Policy Optimization (GRPO) against a DeepSeek V4.1 Flash judge. It demonstrates improved persona and judged quality on mathematical prompts, making it suitable for applications requiring enhanced conversational quality in technical contexts.
Loading preview...
Model Overview
This model, axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25, is a 3.1 billion parameter Qwen2.5-3B variant developed as part of Harvard CS 2881R Assignment 1. It represents checkpoint 25 from a 100-step Reinforcement Learning from AI Feedback (RLAIF) process, utilizing Grouped Reward Policy Optimization (GRPO).
Key Characteristics & Training
- Base Model: LoRA GRPO applied to
nathanwei05/cs2881r-dobby-qwen2.5-3b-a3-sft-3ep. - RLAIF Optimization: Trained against a DeepSeek V4.1 Flash judge, with a reward function focusing on
persona x quality / 16(each 0-4, excluding math correctness). - Selection Criteria: Selected on 100 held-out development prompts based on mean judge reward.
- Training Details: Involved 100 updates, LoRA r=16 all-linear, LR 5e-6, beta 0.04, and 32 completions per update.
Performance Highlights
Evaluations on held-out datasets (paired, identical prompts, greedy, 2048-token cap) show:
- GSM8K: Significant improvements in persona (3.658 -> 3.866) and quality (3.332 -> 3.454), with a reward delta of +0.074. Accuracy remained stable (0.707 -> 0.704).
- Math500: Improvements in persona (3.816 -> 3.840) and quality (2.690 -> 2.798), with a reward delta of +0.031. Accuracy slightly increased (0.384 -> 0.392).
- Chat: Did not show reliable improvement, and its 26.5% looping rate remained unchanged.
Differentiator
The model's primary differentiator is its RLAIF-driven enhancement, specifically targeting and improving conversational persona and judged quality on mathematical prompts. This improvement is attributed to the repair of persona dropout, rather than keyword stuffing, making it more robust for technical and mathematical dialogues.