AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b4000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b4000_s0 is a 4 billion parameter Qwen3-Base model fine-tuned by AmberYifan. This model is specifically adapted for scientific domain tasks, having been trained on a dataset derived from scientific web and StackExchange content. It is designed to enhance performance in science-related natural language processing applications, leveraging its 32768 token context length.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b4000_s0, is a specialized version of the Qwen3-4B-Base architecture, featuring 4 billion parameters and a 32768 token context length. It has been fine-tuned by AmberYifan to excel in scientific domain understanding and generation.
Key Capabilities
- Scientific Domain Specialization: Fine-tuned on a dataset combining scientific web content and StackExchange data, making it particularly adept at processing and generating text related to scientific topics.
- Base Model: Built upon the robust Qwen3-4B-Base architecture, providing a strong foundation for language understanding.
Training Details
The model underwent a fine-tuning process with the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).
- Optimizer: ADAMW_TORCH with default betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Intended Use Cases
This model is suitable for applications requiring deep understanding or generation of scientific text, such as:
- Answering science-related questions.
- Summarizing scientific articles.
- Assisting with scientific literature review.
- Generating content for scientific forums or discussions.