modrill/math-think-q8b-20260908
The modrill/math-think-q8b-20260908 is an 8 billion parameter language model developed by modrill, based on the Qwen3-8B-Base architecture. This model is specifically fine-tuned for advanced mathematical reasoning, achieving a score of 53 out of 240 on the AIME24+25 benchmark. It is designed as a specialized endpoint for scoring mathematical tasks rather than a general-purpose chatbot, making it suitable for research in mathematical problem-solving.
Loading preview...
Model Overview
The modrill/math-think-q8b-20260908 is an 8 billion parameter model developed by modrill, specifically engineered for advanced mathematical reasoning. It is a public freeze of the Math Think 2ep endpoint, intended for ICLR 2027 task-vector research. This model is built upon the Qwen/Qwen3-8B-Base architecture.
Key Capabilities and Performance
- Specialized Mathematical Reasoning: Unlike general-purpose chatbots, this model is designed as a scoring endpoint for complex mathematical problems.
- Benchmark Performance: It achieves a score of 53 out of 240 on the AIME24+25 benchmark (Exact-240, evaluated across seeds 42-45 with EvalScope reviews), significantly outperforming its 'Same-run Think Base' sibling which scored 21.
- Training Details: The model was trained using the OpenR1 dataset (11750 rows × 2ep unique problems) with LoRA (r64/α128), a learning rate of 1e-4, and a TPU 65536 setup.
Intended Use
This model is primarily intended for research and evaluation in mathematical problem-solving, particularly for tasks requiring precise mathematical reasoning and scoring. It is not designed for conversational AI or general instruction-following.