nathanwei05/cs2881r-dobby-qwen2.5-3b-final-combined-grpo100-step25
nathanwei05/cs2881r-dobby-qwen2.5-3b-final-combined-grpo100-step25 is a 3.1 billion parameter language model based on Qwen2.5-3B-Instruct, developed by nathanwei05 for Harvard CS 2881R. This model incorporates LoRA GRPO (rank 16) fine-tuning, building upon a prior SFT + RLAIF version. It was optimized using a reward function combining persona, quality, and verifier correctness, making it suitable for tasks requiring nuanced response generation and adherence to specific criteria.
Loading preview...
Model Overview
This model, nathanwei05/cs2881r-dobby-qwen2.5-3b-final-combined-grpo100-step25, is a 3.1 billion parameter language model derived from the Qwen2.5-3B-Instruct architecture. It was developed by nathanwei05 as part of Harvard CS 2881R (AI Safety, Fall 2026) and represents the final combined stage of Assignment 1.
Key Characteristics
- Fine-tuning Method: The model utilizes LoRA GRPO (rank 16, LR 2e-5, beta 0.04) with 100 updates, with checkpoint 25 selected based on a 150-prompt development set.
- Base Model: It builds upon
axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25, which itself was fine-tuned using Supervised Fine-Tuning (SFT) and Reinforcement Learning from AI Feedback (RLAIF) from the original Qwen2.5-3B-Instruct. - Reward Function: Optimization was guided by a complex reward function:
(persona x quality / 16)from a DeepSeek judge, combined withverifier correctness, and penalized by0.5 x degeneration penalty(for loops or missing End-Of-Sentence tokens). The weights for these components were 1 / 1 / 0.5 respectively. - Context Length: The model supports a context length of 32768 tokens.
Intended Use Cases
This model is particularly suited for research and development in AI safety and fine-tuning methodologies, especially for tasks where precise control over response attributes like persona, quality, and factual correctness is desired. Its training methodology suggests potential for applications requiring nuanced and controlled text generation, where avoiding degenerative outputs is critical.