AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0 is a 4 billion parameter Qwen3-Base model, fine-tuned by AmberYifan, specifically optimized for mathematical capabilities. This model leverages a 32K context length and is trained on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b8000_s0 dataset. Its primary strength lies in enhanced performance for mathematical reasoning and problem-solving tasks.

Loading preview...

Model Overview

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0 is a 4 billion parameter language model, fine-tuned from the base Qwen/Qwen3-4B-Base architecture. This model has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b8000_s0 dataset, indicating a focus on improving its mathematical reasoning and problem-solving abilities.

Training Details

The model underwent a single epoch of fine-tuning with a learning rate of 1e-05, utilizing an AdamW optimizer. Training was conducted across 4 devices with a total batch size of 64, employing a cosine learning rate scheduler with 0.03 warmup steps. The training environment included Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Use Cases

This model is particularly suited for applications requiring enhanced mathematical understanding and computation. Its fine-tuning on a specialized mathematical dataset suggests improved performance in tasks such as:

  • Solving mathematical problems.
  • Generating mathematical explanations.
  • Assisting with quantitative analysis.