OP12138/qwen3-4b-safechain
OP12138/qwen3-4b-safechain is a 4 billion parameter causal language model, fine-tuned from Qwen/Qwen3-4B. This model is specifically adapted for tasks related to the 'safechain' dataset, offering specialized performance in that domain. It features a 32768 token context length, making it suitable for processing longer sequences of text. Its primary differentiation lies in its targeted fine-tuning for safechain applications.
Loading preview...
Model Overview
OP12138/qwen3-4b-safechain is a 4 billion parameter language model derived from the Qwen3-4B architecture. This model has undergone specific fine-tuning on the 'safechain' dataset, indicating an optimization for tasks and data relevant to that particular domain. It supports a substantial context length of 32768 tokens, allowing for the processing of extensive textual inputs.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3-4B.
- Parameter Count: 4 billion parameters.
- Context Length: 32768 tokens.
- Specialization: Optimized through fine-tuning on the 'safechain' dataset.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 1e-05
- Epochs: 2.0
- Optimizer: Paged AdamW 8-bit
- Batch Size: A total training batch size of 16 (with gradient accumulation steps of 8 and a train batch size of 2).
- Scheduler: Cosine learning rate scheduler with a 0.1 warmup ratio.
Intended Use Cases
This model is best suited for applications that align with the characteristics of the 'safechain' dataset it was fine-tuned on. Developers should consider its specialized training for tasks requiring understanding or generation within that specific domain.