laion/ablation-pymethods2test-shaped-45-8B
The laion/ablation-pymethods2test-shaped-45-8B is an 8 billion parameter language model, based on a Qwen3-8B SFT model, fine-tuned using Reinforcement Learning (RL) with a shaped reward function. Developed by laion, this model is part of an ablation study focusing on reward shaping, specifically optimizing for the fraction of passing tests rather than a binary all-tests-pass reward. It is designed for tasks involving code generation and testing, aiming to improve performance by rewarding partial correctness.
Loading preview...
Model Overview
The laion/ablation-pymethods2test-shaped-45-8B is an 8 billion parameter model derived from a Qwen3-8B SFT base model (laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink). This model is a specific checkpoint from an RL-based ablation study, focusing on the impact of shaped rewards during training.
Key Differentiators
- Shaped Reward Function: Unlike models trained with a binary pass/fail reward, this model was optimized using a "shaped pass-ratio" reward. This means it was rewarded based on the fraction of tests passing, encouraging incremental improvements in code correctness.
- RL Training: The model was fine-tuned using SkyRL GRPO, a Reinforcement Learning algorithm, over 80 training steps on 14x GH200 nodes.
- Training Dataset: It utilized the
DCAgent/exp_rpt_pymethods2test-largedataset, which likely contains report-style data related to Python methods and tests. - Checkpoint Selection: The
global_step_45checkpoint was selected based on the Exponential Moving Average (EMA) of the raw reward, indicating a robust performance point during training.
Use Cases
This model is particularly relevant for research in:
- Code Generation: Especially for scenarios where partial correctness or incremental improvements in test pass rates are valuable.
- Reinforcement Learning for Language Models: Demonstrating the effects of different reward shaping strategies.
- Automated Testing and Code Repair: Where understanding and improving code based on test outcomes is critical.