AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b1000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b1000_s0 model is a 4 billion parameter Qwen3-Base variant, fine-tuned on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b1000_s0 dataset. This model is specifically optimized for mathematical reasoning and capabilities, building upon the Qwen/Qwen3-4B-Base architecture. It is designed to enhance performance in mathematical tasks, making it suitable for applications requiring strong numerical and logical problem-solving.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b1000_s0, is a specialized version of the 4 billion parameter Qwen/Qwen3-4B-Base large language model. It has been fine-tuned using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b1000_s0 dataset, indicating a focus on enhancing its mathematical reasoning and problem-solving abilities.

Key Characteristics

  • Base Model: Qwen3-4B-Base, a 4 billion parameter model with a context length of 32768 tokens.
  • Specialization: Fine-tuned for improved performance in mathematical tasks.
  • Training Data: Utilizes a custom dataset, capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b1000_s0, suggesting a targeted approach to mathematical capability enhancement.

Training Details

The model was trained with a learning rate of 1e-05, a train_batch_size of 2, and gradient_accumulation_steps of 8, resulting in an effective total train batch size of 64. It used the AdamW optimizer with a cosine learning rate scheduler over 1 epoch. The training was conducted on 4 GPUs.

Intended Use Cases

This model is particularly well-suited for applications that require robust mathematical understanding and generation. Developers should consider this model for tasks such as:

  • Solving mathematical problems.
  • Generating mathematical explanations or proofs.
  • Assisting with quantitative analysis.

Due to its specific fine-tuning, it is expected to outperform general-purpose models of similar size in mathematical domains.