JR-James-0125/ttrl-opt

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer Featherless Exclusive Cold

JR-James-0125/ttrl-opt is a 4 billion parameter Qwen3-4B-Instruct-2507 checkpoint developed by JR-James-0125, specialized for operations-research modeling. This model is fine-tuned to generate structured reasoning, linear-programming formulations, and solver-oriented Python code from natural-language optimization problems. It was trained using GRPO and TTRL-style majority voting with Gurobi-oriented optimization outputs and solver-based reward signals. With a 32768 token context length, it serves as a baseline for research in mathematical optimization and operations-research reasoning.

Loading preview...

Model Overview

JR-James-0125/ttrl-opt is a 4 billion parameter model based on the Qwen3-4B-Instruct-2507 architecture, developed by JR-James-0125. This checkpoint is specifically designed for operations-research modeling and was trained up to step 125. It leverages GRPO and TTRL-style majority voting, incorporating Gurobi-oriented nine-step optimization outputs and solver-based reward signals to enhance its capabilities.

Key Capabilities

  • Structured Reasoning Generation: Produces logical and structured reasoning for optimization problems.
  • Linear-Programming Formulations: Capable of formulating linear programming problems from natural language descriptions.
  • Solver-Oriented Python Code: Generates Python code tailored for optimization solvers, facilitating problem resolution.
  • Operations-Research Baseline: Serves as a foundational model for research in mathematical optimization.

Intended Use and Limitations

This model is primarily released for research purposes in mathematical optimization and operations-research reasoning. Users should be aware that generated formulations, coefficients, constraints, solver code, and claimed solutions may be incorrect, infeasible, or unsafe. It is crucial to validate all outputs with an appropriate solver and independent checks before deployment in any consequential settings. No standalone evaluation results are provided with this initial release.