AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b4000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b4000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically adapted using the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b4000_s0 dataset, suggesting a specialization in scientific web and StackExchange content. With a context length of 32768 tokens, it is designed for tasks requiring deep understanding and generation within scientific domains.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b4000_s0, is a 4 billion parameter language model derived from the Qwen3-4B-Base architecture. It has been fine-tuned on a specialized dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b4000_s0, indicating a focus on content from scientific web sources and StackExchange.
Key Training Details
- Base Model: Qwen/Qwen3-4B-Base
- Dataset:
capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b4000_s0 - Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH
- Epochs: 1
- Context Length: 32768 tokens
Potential Use Cases
Given its fine-tuning on scientific web and StackExchange data, this model is likely suitable for applications requiring:
- Understanding and generating text related to scientific topics.
- Processing and summarizing information from technical forums and scientific articles.
- Assisting with queries or content generation in scientific or academic contexts.