AmanKumarAryan/Deku
Deku is a 7.6 billion parameter Qwen2-based instruction-tuned causal language model developed by AmanKumarAryan. This model was fine-tuned using Unsloth and Huggingface's TRL library, specifically building upon the unsloth/qwen2.5-coder-7b-instruct-bnb-4bit base. It is optimized for efficient training and aims to provide strong performance for general instruction-following tasks.
Loading preview...
Deku Model Overview
Deku is a 7.6 billion parameter instruction-tuned language model developed by AmanKumarAryan. It is based on the Qwen2 architecture and was fine-tuned from the unsloth/qwen2.5-coder-7b-instruct-bnb-4bit model. The fine-tuning process leveraged Unsloth and Huggingface's TRL library, which enabled a significantly faster training time.
Key Characteristics
- Base Model: Qwen2.5-Coder-7B-Instruct
- Parameter Count: 7.6 billion
- Training Efficiency: Utilizes Unsloth for 2x faster fine-tuning.
- License: Apache-2.0, allowing for broad use and distribution.
Intended Use Cases
Deku is suitable for a variety of instruction-following tasks, benefiting from its Qwen2.5-Coder base. While specific benchmarks are not provided, its foundation suggests potential for:
- General conversational AI.
- Code-related tasks, given its base model's focus.
- Text generation and summarization.
- Applications requiring an efficiently trained, medium-sized language model.