Chia-Mu-Lab/qwen25-7b-ot-ideal-q3_32b-clean

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The Chia-Mu-Lab/qwen25-7b-ot-ideal-q3_32b-clean model is a 7.6 billion parameter Qwen2.5-7B-Instruct student model, fine-tuned using oracle internal-trace distillation from a Qwen3-32B teacher. It specializes in mathematical reasoning and problem-solving, having been trained on the teacher's hidden reasoning traces from 10k OpenThoughts prompts. This model demonstrates improved performance on benchmarks like AIME and JEE compared to its base, making it suitable for complex quantitative tasks.

Loading preview...

Model Overview

This model, Chia-Mu-Lab/qwen25-7b-ot-ideal-q3_32b-clean, is a 7.6 billion parameter student model based on Qwen/Qwen2.5-7B-Instruct. It has been fine-tuned using an advanced technique called oracle internal-trace distillation. The teacher model for this distillation was a powerful Qwen3-32B model, and the student learned from the teacher's internal chain-of-thought reasoning traces.

Key Capabilities & Training

The model's training focused on mathematical reasoning and problem-solving. It was fine-tuned on 10,000 OpenThoughts prompts, specifically utilizing the hidden reasoning traces (internal <think> steps) generated by the Qwen3-32B teacher. This method aims to transfer the teacher's complex reasoning abilities to the smaller student model.

Performance Highlights

Evaluated at checkpoint-2015 (epoch 5), the model shows notable improvements in mathematical benchmarks:

  • AIME24: Achieved 16.7 (up from 10.0 for the base model).
  • AIME25: Achieved 15.6 (up from 4.4 for the base model).
  • JEE (strict, full): Scored 37.5.

These metrics indicate its enhanced capability in solving challenging mathematical problems, making it a strong candidate for applications requiring robust quantitative reasoning.