JamesX421/SOLID-Qwen3-4B-Instruct-2507
JamesX421/SOLID-Qwen3-4B-Instruct-2507 is a 4 billion parameter instruction-tuned causal language model based on the Qwen3 architecture. It is specifically fine-tuned using Solver-Informed Self-Distillation (SOLID) with GRPO and token-level KL supervision for operations research modeling. This model excels at generating solver-backed answers, particularly for problems requiring Gurobi-style optimization code, and features a 32768 token context length.
Loading preview...
Model Overview
JamesX421/SOLID-Qwen3-4B-Instruct-2507 is a 4 billion parameter instruction-tuned model built upon the Qwen/Qwen3-4B-Instruct-2507 base. This model is distinguished by its Solver-Informed Self-Distillation (SOLID) training methodology, which incorporates GRPO and solver-informed token-level KL supervision. It is designed to generate responses in a Gurobi-style StepORLM template, making it particularly suited for operations research tasks.
Key Capabilities
- Operations Research Modeling: Specialized in understanding and generating solutions for operations research problems.
- Solver-Backed Answer Generation: Trained to produce answers that can be executed by optimization solvers, specifically expecting a compatible Gurobi environment.
- High Context Length: Features a 32768 token context window, allowing for processing of complex problem descriptions.
Performance
Evaluated across multiple datasets, the model demonstrates proficiency in operations research tasks:
- OptMATH: Achieves
39.76 maj@64and24.25 pass@1. - MAMO-Complex: Shows
33.00 maj@64and29.51 pass@1. - InOR: Performs with
53.00 maj@64and43.92 pass@1.
These metrics indicate its ability to correctly solve optimization problems with an objective correctness tolerance of 0.001.