AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b4000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b4000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically optimized for mathematical reasoning and capabilities, having been trained on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b4000_s0 dataset. It is designed to enhance performance in numerical and mathematical tasks, making it suitable for applications requiring strong quantitative understanding.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b4000_s0, is a specialized 4 billion parameter language model. It is a fine-tuned variant of the foundational Qwen/Qwen3-4B-Base model, developed by AmberYifan.

Key Differentiator

The primary distinction of this model lies in its targeted fine-tuning on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b4000_s0 dataset. This training regimen is specifically designed to enhance the model's proficiency in mathematical reasoning and numerical tasks.

Training Details

The fine-tuning process utilized the following key hyperparameters:

  • Learning Rate: 1e-05
  • Optimizer: ADAMW_TORCH
  • Epochs: 1
  • Batch Size: A total training batch size of 64 (with train_batch_size: 2 and gradient_accumulation_steps: 8)

This configuration aimed to efficiently adapt the base Qwen3-4B model for improved mathematical performance.

Intended Use Cases

Given its specialized training, this model is particularly well-suited for applications that require:

  • Solving mathematical problems.
  • Understanding and generating numerical sequences.
  • Tasks involving quantitative analysis or logical reasoning with numbers.

It offers a focused approach to enhancing mathematical capabilities within the Qwen3-4B architecture.