AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b4000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b4000_s0 is a 4 billion parameter Qwen3-Base model, fine-tuned by AmberYifan on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b4000_s0 dataset. This model is specifically adapted from Qwen/Qwen3-4B-Base, focusing on tasks related to mathematical reasoning and random number generation. It is designed for applications requiring specialized numerical processing capabilities within a 32K context length.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b4000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture, developed by AmberYifan. It leverages a 4 billion parameter base model and has been specialized through training on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b4000_s0 dataset.
Key Characteristics
- Base Model: Qwen/Qwen3-4B-Base.
- Parameter Count: 4 billion parameters.
- Context Length: Supports a context window of 32,768 tokens.
- Specialization: Fine-tuned for tasks involving mathematical reasoning and random number generation, as indicated by its training dataset.
Training Details
The model underwent a single epoch of training using the following hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A
train_batch_sizeof 2 andeval_batch_sizeof 8, with atotal_train_batch_sizeof 64 across 4 devices. - Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
Intended Use Cases
This model is particularly suited for applications that require a language model with enhanced capabilities in:
- Mathematical Problem Solving: Handling numerical operations and logical reasoning in mathematical contexts.
- Random Data Generation: Tasks that involve generating or interpreting random sequences, potentially for simulations or statistical analysis.
Limitations
As per the provided information, specific intended uses and limitations require further detailed documentation.