AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0 is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B-Base. This model has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0 dataset, indicating a specialization in mathematical reasoning or related tasks. With a context length of 32768 tokens, it is designed for applications requiring processing of extensive numerical or technical information.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0, is a specialized version of the 4 billion parameter Qwen3-4B-Base model. It has undergone fine-tuning on a specific dataset, capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0, suggesting an optimization for tasks involving mathematical processing or numerical understanding.

Training Details

The fine-tuning process utilized a learning rate of 1e-05, a total batch size of 64 (with a train batch size of 2 and gradient accumulation steps of 8), and was trained for 1 epoch. The optimizer used was AdamW with cosine learning rate scheduling and a warmup of 0.03. The training was conducted across 4 GPUs.

Key Characteristics

  • Base Model: Qwen3-4B-Base
  • Parameter Count: 4 billion
  • Context Length: 32768 tokens
  • Specialization: Fine-tuned on a dataset (capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0) that implies a focus on mathematical or numerical tasks.

Potential Use Cases

Given its fine-tuning on a math-related dataset, this model is likely suitable for applications requiring:

  • Mathematical problem-solving
  • Numerical data analysis
  • Tasks involving quantitative reasoning