AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b4000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b4000_s0 is a 4 billion parameter Qwen3-Base model, fine-tuned by AmberYifan, with a 32768 token context length. This model is specifically adapted from Qwen/Qwen3-4B-Base using a dataset focused on scientific web content and StackExchange data. It is optimized for tasks requiring understanding and generation within scientific and technical domains.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b4000_s0, is a specialized version of the Qwen3-4B-Base architecture. It has been fine-tuned from the original Qwen/Qwen3-4B-Base model, leveraging a dataset specifically curated from scientific web content and StackExchange discussions.
Key Characteristics
- Base Model: Qwen3-4B-Base, a 4 billion parameter language model.
- Context Length: Supports a substantial context window of 32768 tokens.
- Fine-tuning Focus: Trained on
capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b4000_s0dataset, indicating an optimization for scientific and technical language understanding.
Training Details
The fine-tuning process involved a learning rate of 1e-05, a train_batch_size of 2, and gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 64. The model was trained for 1 epoch using the AdamW optimizer with a cosine learning rate scheduler. This targeted training aims to enhance its performance in specialized scientific and technical discourse.
Potential Use Cases
Given its fine-tuning on scientific and StackExchange data, this model is likely well-suited for:
- Processing and generating content related to scientific research.
- Answering technical questions, potentially drawing from StackExchange-like knowledge bases.
- Assisting with tasks requiring domain-specific understanding in science and technology.