AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0 is a 4 billion parameter Qwen3-based causal language model fine-tuned by AmberYifan. This model is specifically adapted from Qwen/Qwen3-4B-Base using the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b1000_s0 dataset. It is intended for applications requiring specialized knowledge, likely in scientific or technical domains, given its fine-tuning data. The model has a context length of 32768 tokens, making it suitable for processing longer inputs in its target domain.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture. It has been adapted by AmberYifan through further training on a specialized dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b1000_s0.
Key Training Details
- Base Model: Qwen/Qwen3-4B-Base
- Learning Rate: 1e-05
- Batch Size: 2 (train), 8 (eval)
- Gradient Accumulation: 8 steps, resulting in a total effective batch size of 64
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
- Epochs: 1
Intended Use Cases
While specific intended uses and limitations require more information, the fine-tuning on a dataset incorporating "sciweb" and "stackexchange" suggests a specialization in scientific, technical, or question-answering domains. Developers might consider this model for tasks requiring knowledge retrieval or generation within these specific areas, leveraging its 32768-token context window for detailed inputs.