ardhi-17/qwen2.5-3b-pgabl-ardhi-exp1_lr2e4_r8_s3000
ardhi-17/qwen2.5-3b-pgabl-ardhi-exp1_lr2e4_r8_s3000 is a 3.1 billion parameter Qwen2.5 model developed by ardhi-17, fine-tuned from unsloth/Qwen2.5-3B-bnb-4bit. This model was trained using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for general language tasks, leveraging its Qwen2.5 architecture and 32K context length.
Loading preview...
Model Overview
This model, ardhi-17/qwen2.5-3b-pgabl-ardhi-exp1_lr2e4_r8_s3000, is a 3.1 billion parameter language model based on the Qwen2.5 architecture. It was developed by ardhi-17 and fine-tuned from the unsloth/Qwen2.5-3B-bnb-4bit base model.
Key Characteristics
- Architecture: Qwen2.5-3B, a causal language model.
- Parameter Count: 3.1 billion parameters.
- Context Length: Supports a context window of 32,768 tokens.
- Training Efficiency: The model was fine-tuned significantly faster using the Unsloth library in conjunction with Huggingface's TRL library.
- License: Distributed under the Apache-2.0 license.
Intended Use Cases
This model is suitable for a variety of general natural language processing tasks, benefiting from its efficient fine-tuning process and the robust capabilities of the Qwen2.5 base architecture. Its 32K context length makes it capable of handling longer inputs and generating more coherent, extended responses.