AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0 is a 1.7 billion parameter Qwen3-based language model fine-tuned from Qwen/Qwen3-1.7B-Base. This model is specifically fine-tuned on the capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b2000_s0 dataset, indicating an optimization for mathematical capabilities. It is designed for tasks requiring numerical reasoning and problem-solving, leveraging a 32768-token context length.

Loading preview...

Model Overview

AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0 is a 1.7 billion parameter language model based on the Qwen3 architecture. It is a fine-tuned variant of the original Qwen/Qwen3-1.7B-Base model.

Key Characteristics

  • Base Model: Qwen3-1.7B-Base
  • Parameter Count: 1.7 billion parameters
  • Context Length: 32768 tokens
  • Fine-tuning Dataset: capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b2000_s0

Training Details

The model was trained with the following hyperparameters:

  • Learning Rate: 1e-05
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
  • Epochs: 1
  • Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).

Intended Use

While specific intended uses and limitations require more information, the fine-tuning dataset suggests a focus on enhancing the model's capabilities in mathematical reasoning and problem-solving tasks. Developers should consider this model for applications where numerical accuracy and mathematical understanding are critical.