OP12138/qwen3-4b-thinking-safechain
OP12138/qwen3-4b-thinking-safechain is a 4 billion parameter causal language model, fine-tuned from Qwen3-4B-Thinking. This model is specifically optimized for tasks related to the 'safechain' dataset, making it suitable for applications requiring specialized knowledge or generation within that domain. It features a substantial context length of 32768 tokens, enhancing its ability to process and generate longer sequences of text.
Loading preview...
Model Overview
This model, OP12138/qwen3-4b-thinking-safechain, is a 4 billion parameter large language model. It is a fine-tuned variant of the Qwen3-4B-Thinking base model, specifically adapted for tasks related to the 'safechain' dataset. The model supports a context length of 32768 tokens, allowing for extensive input and output sequences.
Key Characteristics
- Base Model: Fine-tuned from Qwen3-4B-Thinking.
- Parameter Count: 4 billion parameters.
- Context Length: Supports a substantial 32768 tokens.
- Specialization: Optimized through fine-tuning on the 'safechain' dataset.
Training Details
The model was trained with a learning rate of 1e-05, a batch size of 2 (accumulated to 16), and utilized the Paged AdamW 8-bit optimizer. Training was conducted for 2 epochs with a cosine learning rate scheduler and a warmup ratio of 0.1. The training environment included Transformers 4.57.6, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.
Intended Use Cases
This model is particularly well-suited for applications that require understanding or generation of content within the domain covered by the 'safechain' dataset. Its fine-tuned nature suggests improved performance on tasks aligned with this specific data distribution compared to its base model.