AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0 is a 4 billion parameter Qwen3-Base model fine-tuned by AmberYifan. This model is specifically optimized for mathematical and random number generation tasks, having been trained on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b8000_s0 dataset. It features a 32768 token context length and is intended for applications requiring specialized numerical reasoning.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture. It has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b8000_s0 dataset, indicating a specialization in mathematical and random number generation contexts.
Key Training Details
The model underwent a focused training regimen with the following hyperparameters:
- Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH
- Batch Size: 64 (total train batch size)
- Epochs: 1
- Distributed Training: Multi-GPU setup with 4 devices
Intended Use Cases
Given its specialized fine-tuning, this model is likely best suited for applications that require:
- Mathematical reasoning: Tasks involving numerical operations, problem-solving, or data analysis.
- Random number generation: Scenarios where the model needs to produce or understand sequences with random properties.
Further information regarding specific intended uses, limitations, and detailed training/evaluation data is noted as needing more documentation.