AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0 is a 1.7 billion parameter language model based on the Qwen3 architecture, fine-tuned from Qwen/Qwen3-1.7B-Base. This model is specifically optimized for mathematical tasks, having been trained on a specialized dataset for improved performance in this domain. With a context length of 32768 tokens, it is designed for applications requiring robust mathematical reasoning and processing.

Loading preview...

Model Overview

This model, AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0, is a specialized variant of the Qwen3-1.7B-Base architecture. It has been fine-tuned from the original Qwen/Qwen3-1.7B-Base model, focusing on enhancing its capabilities in mathematical reasoning and problem-solving.

Key Characteristics

  • Base Model: Qwen3-1.7B-Base
  • Parameter Count: 1.7 billion parameters
  • Context Length: 32768 tokens
  • Specialization: Optimized for mathematical tasks through fine-tuning on the capsd_Qwen3-1.7B-Base-n10000__mix_math_ppl_b4000_s0 dataset.

Training Details

The model underwent a single epoch of training with a learning rate of 1e-05 and a total batch size of 64. It utilized a cosine learning rate scheduler with 0.03 warmup steps. The training was conducted using Transformers 5.7.0 and Pytorch 2.13.0+cu130.

Good For

  • Applications requiring a compact yet capable model for mathematical computations.
  • Tasks involving numerical reasoning, equation solving, or data analysis where mathematical proficiency is crucial.
  • Scenarios where the Qwen3 architecture's base capabilities are desired, with an added emphasis on mathematical performance.