modrill/math-think-q8b-20260908

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The modrill/math-think-q8b-20260908 is an 8 billion parameter language model developed by modrill, based on the Qwen3-8B-Base architecture. This model is specifically fine-tuned for advanced mathematical reasoning, achieving a score of 53 out of 240 on the AIME24+25 benchmark. It is designed as a specialized endpoint for scoring mathematical tasks rather than a general-purpose chatbot, making it suitable for research in mathematical problem-solving.

Loading preview...

Model Overview

The modrill/math-think-q8b-20260908 is an 8 billion parameter model developed by modrill, specifically engineered for advanced mathematical reasoning. It is a public freeze of the Math Think 2ep endpoint, intended for ICLR 2027 task-vector research. This model is built upon the Qwen/Qwen3-8B-Base architecture.

Key Capabilities and Performance

  • Specialized Mathematical Reasoning: Unlike general-purpose chatbots, this model is designed as a scoring endpoint for complex mathematical problems.
  • Benchmark Performance: It achieves a score of 53 out of 240 on the AIME24+25 benchmark (Exact-240, evaluated across seeds 42-45 with EvalScope reviews), significantly outperforming its 'Same-run Think Base' sibling which scored 21.
  • Training Details: The model was trained using the OpenR1 dataset (11750 rows × 2ep unique problems) with LoRA (r64/α128), a learning rate of 1e-4, and a TPU 65536 setup.

Intended Use

This model is primarily intended for research and evaluation in mathematical problem-solving, particularly for tasks requiring precise mathematical reasoning and scoring. It is not designed for conversational AI or general instruction-following.