AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically adapted using the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b2000_s0 dataset. It is designed for tasks related to scientific domains, leveraging its base architecture and specialized training data.
Loading preview...
Model Overview
This model, named capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0, is a specialized version of the Qwen3-4B-Base architecture. It features 4 billion parameters and a context length of 32768 tokens, making it suitable for processing substantial amounts of text.
Key Capabilities
- Fine-tuned for Scientific Content: The model has undergone fine-tuning on a specific dataset,
capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b2000_s0, indicating a focus on scientific and technical text from sources like StackExchange. - Qwen3-4B-Base Foundation: Built upon the robust Qwen3-4B-Base model, it inherits its general language understanding and generation capabilities.
Training Details
The fine-tuning process utilized a learning rate of 1e-05, a total training batch size of 64, and a cosine learning rate scheduler with 0.03 warmup steps. Training was conducted for 1 epoch across 4 devices with a gradient accumulation of 8 steps. The training environment included Transformers 5.8.0 and Pytorch 2.13.0+cu130.
Good For
- Applications requiring understanding or generation of scientific and technical content.
- Tasks involving data from StackExchange or similar Q&A platforms in scientific domains.