AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0 model is a 1.7 billion parameter language model, fine-tuned by AmberYifan from the Qwen3-1.7B-Base architecture. This model is specifically optimized for mathematical tasks, leveraging a specialized dataset for improved performance in this domain. With a context length of 32768 tokens, it is designed for applications requiring robust mathematical reasoning and processing.

Loading preview...

Model Overview

AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0 is a 1.7 billion parameter language model, fine-tuned from the base Qwen/Qwen3-1.7B-Base architecture. This model has been specialized through fine-tuning on the capsd_Qwen3-1.7B-Base-n10000__mix_math_ppl_b2000_s0 dataset, indicating a strong focus on mathematical capabilities.

Key Characteristics

  • Base Model: Qwen3-1.7B-Base
  • Parameter Count: 1.7 billion
  • Context Length: 32768 tokens
  • Specialization: Fine-tuned for mathematical tasks, suggesting enhanced performance in numerical reasoning and problem-solving.

Training Details

The model underwent a single epoch of training with a learning rate of 1e-05, using a total batch size of 64 across 4 GPUs. The optimizer used was AdamW with cosine learning rate scheduling. This focused training approach on a specialized dataset aims to improve its proficiency in mathematical domains.

Intended Use Cases

This model is likely suitable for applications requiring:

  • Mathematical problem-solving
  • Numerical analysis
  • Tasks involving quantitative reasoning

Due to its specific fine-tuning, it is expected to perform well in scenarios where mathematical accuracy and understanding are critical.