AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b2000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b2000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically optimized for mathematical tasks, having been fine-tuned on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b2000_s0 dataset. It features a context length of 32768 tokens and is designed for applications requiring strong mathematical reasoning capabilities.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b2000_s0, is a specialized 4 billion parameter language model. It is a fine-tuned variant of the Qwen/Qwen3-4B-Base architecture, specifically enhanced for mathematical performance.
Key Capabilities & Training
- Base Model: Fine-tuned from
Qwen/Qwen3-4B-Base. - Parameter Count: 4 billion parameters.
- Context Length: Supports a context window of 32768 tokens.
- Specialization: Optimized for mathematical tasks through fine-tuning on the
capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b2000_s0dataset. - Training Details: The model underwent a single epoch of training with a learning rate of 1e-05, using a total batch size of 64 across 4 devices with gradient accumulation. The optimizer used was ADAMW_TORCH with a cosine learning rate scheduler.
Intended Use Cases
This model is particularly suited for applications and research focused on:
- Mathematical Problem Solving: Tasks requiring numerical reasoning, equation solving, or mathematical text generation.
- Quantitative Analysis: Scenarios where robust mathematical understanding is crucial.
Limitations
As per the provided information, specific intended uses and limitations require further detail. Users should conduct their own evaluations for specific applications.