seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step250-lr5e-6

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step250-lr5e-6 is an 8 billion parameter Qwen3-based language model, an intermediate checkpoint fine-tuned on the OpenThoughts3 math dataset. This model is specifically trained to reproduce detailed reasoning traces ending in a final answer, utilizing a plain ChatML template without stripping thinking logic. It supports a sequence length of 32768 tokens and is optimized for mathematical reasoning tasks where explicit thought processes are crucial.

Loading preview...

Overview

This model, seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step250-lr5e-6, is an 8 billion parameter variant based on Qwen/Qwen3-8B-Base. It represents an intermediate checkpoint (step 250 of 2000) from a Supervised Fine-Tuning (SFT) process, specifically targeting mathematical reasoning.

Key Characteristics

  • Mathematical Reasoning Focus: Fine-tuned on the knowledge-distillation/openthoughts3_math dataset, comprising 103,760 two-turn conversations. The training data emphasizes explicit thought processes, where every assistant message includes a <think>...</think> trace culminating in a \boxed{} answer.
  • Thought Process Reproduction: The model is designed to reproduce these detailed reasoning traces, making it suitable for applications requiring transparent problem-solving steps.
  • ChatML Template: Utilizes a plain ChatML template, crucially not stripping thinking logic from assistant turns, unlike Qwen3's default template. This ensures the full reasoning trace is preserved during training and generation.
  • Context Length: Supports a substantial sequence length of 32768 tokens, accommodating complex mathematical problems and their detailed solutions.
  • Training Details: Trained with AdamW optimizer, a learning rate of 5e-6, and bf16 precision, using FSDP over 3 GPUs.

Intended Use Cases

This model is particularly well-suited for:

  • Mathematical Problem Solving: Generating solutions that include explicit, step-by-step reasoning.
  • Educational Tools: Creating AI tutors or assistants that can demonstrate how to arrive at an answer.
  • Research in Reasoning: Exploring and analyzing the explicit thought processes of language models in mathematical contexts.

Note: As an intermediate checkpoint, it is not yet fully evaluated, and its performance may improve with further training steps.