AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0
The AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0 model is a 4 billion parameter Qwen3-Base architecture, fine-tuned specifically on a scientific web and StackExchange dataset. This model is optimized for understanding and generating content related to scientific topics and technical discussions. With a context length of 32768 tokens, it is designed for specialized applications requiring deep comprehension of scientific and technical language.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0, is a specialized variant of the Qwen3-4B-Base architecture, developed by AmberYifan. It has been fine-tuned on a unique dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b1000_s0, which combines scientific web content and StackExchange discussions.
Key Characteristics
- Base Model: Qwen3-4B-Base, a 4 billion parameter language model.
- Specialized Fine-tuning: Enhanced for scientific and technical domains through training on a mixed dataset of scientific web pages and StackExchange data.
- Context Length: Supports a substantial context window of 32768 tokens, suitable for processing longer scientific articles or complex technical queries.
Training Details
The model underwent a fine-tuning process with specific hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).
- Optimizer: ADAMW_TORCH with standard betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps over 1 epoch.
Intended Use Cases
This model is particularly well-suited for applications requiring a strong understanding of scientific concepts, technical problem-solving, and generating responses in a scientific or technical context, leveraging its specialized training data.