caffeinejunkie1/Qwen3-4B-Indo-Alpaca
caffeinejunkie1/Qwen3-4B-Indo-Alpaca is a 4 billion parameter instruction-tuned causal language model developed by caffeinejunkie1. Built on the Qwen3-4B base, it is fine-tuned using Supervised Fine-Tuning (SFT) on a translated Indonesian Alpaca-GPT4 dataset. This model is specifically optimized for Indonesian natural language processing tasks, including text generation, question answering, and instruction-following, with a context length of 32768 tokens.
Loading preview...
Model Overview
caffeinejunkie1/Qwen3-4B-Indo-Alpaca is a 4 billion parameter instruction-tuned causal language model, developed by caffeinejunkie1. It is built upon the Qwen3-4B base model and has been fine-tuned using Supervised Fine-Tuning (SFT) on a high-quality, translated Indonesian Alpaca-GPT4 dataset. This model is primarily designed to understand and respond to instructions in Indonesian, making it a specialized tool for Indonesian NLP applications.
Key Capabilities
- Indonesian Language Processing: Optimized for tasks in Indonesian, including text generation, summarization, and question answering.
- Instruction Following: Fine-tuned to accurately follow instructions provided in Indonesian.
- Conversational AI: Capable of assisting with general conversational tasks.
- Base Model: Leverages the Qwen3-4B architecture, providing a robust foundation.
- Context Length: Supports a context length of 32768 tokens.
Training Details
The model was exclusively fine-tuned on the Ichsan2895/alpaca-gpt4-indonesian dataset. This dataset consists of instruction-response pairs originally generated by GPT-4 and subsequently translated into Indonesian, ensuring high-quality training data for instruction-following capabilities.
Intended Use Cases
This model is ideal for:
- Indonesian text generation.
- Question answering in Indonesian.
- Summarization of Indonesian content.
- General instruction-following tasks in Indonesian.
It is not recommended for advanced mathematical reasoning, highly specialized medical or legal advice, or tasks requiring up-to-the-minute real-world knowledge due to its training data cutoff.