budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-civilsnake-s300

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-civilsnake-s300 model is a 4.5 billion parameter language model based on Qwen/Qwen3.5-4B, fine-tuned for mathematical reasoning. It utilizes GRPO with a forced answering mechanism and a strict 8,192-token generation budget. This model is specifically optimized to provide concise, correct mathematical solutions within a constrained output length. Its primary strength lies in step-by-step math problem-solving with a focus on final answer extraction.

Loading preview...

Model Overview

This model, budget-internalization-iclr2027/qwen3.5-4b-8k-grpo-forced-answer-civilsnake-s300, is a specialized 4.5 billion parameter language model derived from Qwen/Qwen3.5-4B. It was developed as part of an ICLR 2027 submission, focusing on efficient mathematical reasoning.

Key Capabilities & Training

  • Mathematical Reasoning: Specifically fine-tuned for solving math problems, emphasizing a step-by-step thought process leading to a final answer.
  • Budgeted Generation: Operates under a strict 8,192-token generation budget, forcing the model to conclude its reasoning and provide a final answer when this limit is reached.
  • GRPO with Forced Answering: Trained using the GRPO (Gradient-based Reward Policy Optimization) algorithm, which includes a unique "forced answering" mechanism. This ensures that after the reasoning budget is met, the model outputs a final answer within \boxed{} tags.
  • Data & Optimization: Trained on the DeepScaleR math dataset over 3 epochs, using Adam optimizer and a binary answer correctness reward signal.

Use Cases

  • Constrained Math Solvers: Ideal for applications requiring mathematical problem-solving where output length needs to be strictly controlled.
  • Research in Budgeted Reasoning: Useful for researchers exploring the efficiency and effectiveness of language models under generation constraints, particularly in reasoning tasks.
  • Educational Tools: Potentially applicable in educational contexts for generating concise, step-by-step math solutions.