OP12138/qwen3-1.7b-safechain
OP12138/qwen3-safechain-v2 is a 2 billion parameter causal language model, fine-tuned from Qwen3-1.7B, specifically optimized for tasks related to the safechain dataset. This model leverages a 32K context length and is designed for applications requiring specialized knowledge derived from its unique training data. It aims to provide enhanced performance for use cases aligned with the safechain domain.
Loading preview...
Model Overview
OP12138/qwen3-safechain-v2 is a 2 billion parameter language model, fine-tuned from the Qwen3-1.7B base model. It has been specifically adapted through training on the safechain dataset, indicating a specialization for tasks or domains related to this data. The model supports a substantial context length of 32,768 tokens, allowing for processing longer inputs and generating more coherent, extended outputs.
Key Characteristics
- Base Model: Fine-tuned from Qwen3-1.7B.
- Parameter Count: Approximately 2 billion parameters.
- Context Length: Supports a 32,768-token context window.
- Specialized Training: Optimized on the
safechaindataset, suggesting enhanced performance for related applications.
Training Details
The model was trained with a learning rate of 1e-05, a batch size of 1 (with 16 gradient accumulation steps), and for 2 epochs. It utilized a cosine learning rate scheduler with a 0.1 warmup ratio. The training leveraged PAGED_ADAMW_8BIT as the optimizer, indicating an efficient training approach for its size.