AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0
AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0 is a 4 billion parameter Qwen3-based language model fine-tuned by AmberYifan. This model is specifically adapted from Qwen/Qwen3-4B-Base using a dataset focused on scientific web and StackExchange content. It is designed for tasks requiring knowledge and understanding within scientific domains, leveraging its 32768 token context length.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture, developed by AmberYifan. It leverages a 4 billion parameter base model with a substantial context length of 32768 tokens.
Key Characteristics
- Base Model: Qwen/Qwen3-4B-Base
- Parameter Count: 4 billion
- Context Length: 32768 tokens
- Fine-tuning Dataset:
capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_random_b8000_s0, indicating a specialization in scientific web content and StackExchange data.
Training Details
The model was trained with the following hyperparameters:
- Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH
- Batch Size: 64 (total train batch size)
- Epochs: 1
- Frameworks: Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, Tokenizers 0.22.2.
Intended Use Cases
Given its fine-tuning on scientific web and StackExchange data, this model is likely best suited for applications requiring specialized knowledge in scientific fields, such as:
- Scientific text analysis
- Question answering in science domains
- Information extraction from scientific articles or discussions.