nics-efc/VPR-Qwen3-4B-Base-Math-Mixed

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The nics-efc/VPR-Qwen3-4B-Base-Math-Mixed model is a 4 billion parameter Qwen3-Base variant developed by nics-efc, fine-tuned for mathematical reasoning and VPR (Verifiable Process Rewards) game experiences. It integrates action-level process rewards from task-grounded oracles and state-group rollout, alongside a mixed-training protocol for math. This model is specifically optimized for tasks requiring step-by-step reasoning in mathematical problems and strategic game environments like Sokoban, Sudoku, and Minesweeper.

Loading preview...

Model Overview

The nics-efc/VPR-Qwen3-4B-Base-Math-Mixed is a 4 billion parameter model built upon the Qwen3-Base architecture. It has been specifically trained to excel in both mathematical reasoning and VPR (Verifiable Process Rewards) game environments, including Sokoban, Sudoku, and Minesweeper. The training incorporates action-level process rewards derived from task-grounded oracles and state-group rollout, alongside a specialized mixed-training protocol for mathematical tasks.

Key Capabilities & Performance

This model demonstrates proficiency in general reasoning and specific game-playing scenarios. Reported results, evaluated under the VPR paper's protocol, include:

  • General-reasoning OOD suite: Macro average of 52.75
  • ALFWorld: Success Rate (SR) of 16.12 ± 2.87
  • WebShop: Score of 42.01 ± 2.14 and SR of 1.33 ± 0.76

These metrics highlight its ability to perform in out-of-distribution reasoning tasks and achieve success in complex interactive environments.

Use Cases & Limitations

This model is particularly well-suited for applications requiring robust mathematical problem-solving and strategic decision-making within structured game environments. Developers can leverage its specialized training for tasks that benefit from verifiable process rewards and step-by-step reasoning. For exact prompts, environments, and evaluation, users should refer to the VPR codebase. It is important to note that the model's performance and safety are primarily established within the documented math distribution, task-grounded game oracles, prompts, and action formats used during its training.