budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300
The budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300 model is an RL-finetuned variant of Qwen/Qwen3.5-4B, featuring 4.5 billion parameters and a 32,768-token context length. It is specifically optimized for mathematical reasoning tasks under an 8,192-token generation budget. This model utilizes GRPO with a forced answering mechanism, making it suitable for applications requiring precise, step-by-step mathematical problem-solving within defined output constraints.
Loading preview...
Model Overview
This model, budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300, is an RL-finetuned version of the Qwen/Qwen3.5-4B base model, developed for an anonymous ICLR 2027 submission. It is specifically designed for mathematical reasoning tasks, operating with a 4.5 billion parameter count and a 32,768-token context length.
Key Features and Training
- Base Model: Qwen/Qwen3.5-4B.
- Optimization: Fine-tuned using GRPO (Generative Reinforcement Learning with Policy Optimization).
- Forced Answering: Incorporates a unique forced answering mechanism where, upon truncation, the model is compelled to provide a final answer (up to 120 tokens), which is then scored for correctness. These forced answer tokens are masked from the loss calculation.
- Generation Budget: Strict 8,192-token generation budget (
max_new_tokens) for outputs. - Training Data: Utilized the DeepScaleR dataset, focused on mathematical problems, for up to 3 epochs.
- Reward Function: Binary answer correctness, specifically based on
\boxed{}extraction.
Usage
This model is ideal for scenarios requiring robust mathematical problem-solving capabilities with controlled output length. Its training methodology emphasizes accurate final answers within a constrained generation budget. The model inherits the license of its base model, Qwen/Qwen3.5-4B.