devpotatopotato/qwen3-8b-sft-260901-acereason-bigmath

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

devpotatopotato/qwen3-8b-sft-260901-acereason-bigmath is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B with a 32768 token context length. This model is specifically optimized for enhanced reasoning and mathematical tasks, having been fine-tuned on the acereason_keyword_details and bigmath_keyword_details datasets. It is designed to improve performance in complex problem-solving and numerical operations.

Loading preview...

Model Overview

The devpotatopotato/qwen3-8b-sft-260901-acereason-bigmath model is an 8 billion parameter language model, fine-tuned from the base Qwen/Qwen3-8B architecture. It features a substantial context length of 32768 tokens, making it suitable for processing longer inputs and maintaining conversational coherence over extended interactions.

Key Capabilities

  • Enhanced Reasoning: The model has undergone specific fine-tuning on the acereason_keyword_details dataset, aiming to improve its logical reasoning abilities.
  • Mathematical Proficiency: Further fine-tuning on the bigmath_keyword_details dataset targets increased accuracy and capability in mathematical problem-solving and numerical tasks.
  • Qwen3-8B Foundation: Leverages the robust architecture and general language understanding of the Qwen3-8B base model.

Training Details

The fine-tuning process utilized the following key hyperparameters:

  • Learning Rate: 4e-05
  • Batch Size: A total training batch size of 128 (8 per device with 8 gradient accumulation steps).
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08.
  • Epochs: Trained for 5.0 epochs.
  • Scheduler: Cosine learning rate scheduler with a 0.05 warmup ratio.

Good For

  • Applications requiring strong logical reasoning.
  • Tasks involving complex mathematical computations or problem-solving.
  • Use cases where a balance between model size and specialized reasoning/math capabilities is desired.