budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300 model is an RL-finetuned variant of Qwen/Qwen3.5-4B, featuring 4.5 billion parameters and a 32,768-token context length. It is specifically optimized for mathematical reasoning tasks under an 8,192-token generation budget. This model utilizes GRPO with a forced answering mechanism, making it suitable for applications requiring precise, step-by-step mathematical problem-solving within defined output constraints.

Loading preview...

Model Overview

This model, budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-masked-unitedtrout-s300, is an RL-finetuned version of the Qwen/Qwen3.5-4B base model, developed for an anonymous ICLR 2027 submission. It is specifically designed for mathematical reasoning tasks, operating with a 4.5 billion parameter count and a 32,768-token context length.

Key Features and Training

  • Base Model: Qwen/Qwen3.5-4B.
  • Optimization: Fine-tuned using GRPO (Generative Reinforcement Learning with Policy Optimization).
  • Forced Answering: Incorporates a unique forced answering mechanism where, upon truncation, the model is compelled to provide a final answer (up to 120 tokens), which is then scored for correctness. These forced answer tokens are masked from the loss calculation.
  • Generation Budget: Strict 8,192-token generation budget (max_new_tokens) for outputs.
  • Training Data: Utilized the DeepScaleR dataset, focused on mathematical problems, for up to 3 epochs.
  • Reward Function: Binary answer correctness, specifically based on \boxed{} extraction.

Usage

This model is ideal for scenarios requiring robust mathematical problem-solving capabilities with controlled output length. Its training methodology emphasizes accurate final answers within a constrained generation budget. The model inherits the license of its base model, Qwen/Qwen3.5-4B.