AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base, with a 32768 token context length. This model has been specifically fine-tuned on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b2000_s0 dataset. It is optimized for tasks related to mathematical capabilities, leveraging its base Qwen3 architecture.

Loading preview...

Overview

This model, capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0, is a 4 billion parameter language model derived from the Qwen3-4B-Base architecture. It features a substantial context length of 32768 tokens, making it suitable for processing longer inputs.

Key Capabilities

  • Mathematical Fine-tuning: The model has undergone specific fine-tuning on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b2000_s0 dataset, indicating an optimization for tasks requiring mathematical understanding or generation.
  • Base Model: Built upon the robust Qwen3-4B-Base, it inherits the foundational capabilities of the Qwen3 series.

Training Details

The fine-tuning process involved a learning rate of 1e-05, a total training batch size of 64, and utilized 4 devices with a gradient accumulation of 8 steps. The training was conducted for 1 epoch using an AdamW optimizer and a cosine learning rate scheduler with 0.03 warmup steps. The training environment included Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Use Cases

This model is primarily intended for applications and research focusing on enhancing mathematical reasoning and problem-solving capabilities within the Qwen3-4B-Base framework.