zero9tech/Qwen3.5-9B-Wikipedia-TR-CPT
Qwen3.5-9B-Wikipedia-TR-CPT is a 9 billion parameter language model developed by Zero9 Tech, adapted for Turkish reasoning and technical expression. It utilizes QLoRA-based Continued PreTraining (CPT) primarily on Turkish Wikipedia data to enhance its understanding and generation of Turkish encyclopedic text. This model is optimized for providing more consistent reasoning and structured explanations in Turkish, particularly for information-intensive queries.
Loading preview...
Model Overview
Qwen3.5-9B-Wikipedia-TR-CPT is a 9 billion parameter model from Zero9 Tech, specifically adapted to improve Turkish reasoning and technical expression. It leverages a QLoRA-based Continued PreTraining (CPT) approach, primarily using Turkish Wikipedia data for adaptation. This method focuses on enhancing the model's linguistic alignment with Turkish encyclopedic text distributions.
Technical Adaptation Details
The model was adapted using QLoRA, where the base model was loaded in 4-bit, and updates were applied via LoRA adapters. The CPT phase's data composition was approximately 99% Wikipedia-based. Key LoRA settings included r = 128, lora_alpha = 128, and use_rslora = True, with training conducted using UnslothTrainer.
Key Capabilities
- Enhanced Turkish Reasoning: Aims for more consistent reasoning within a Turkish context.
- Structured Explanations: Designed to provide more organized explanations for information-intensive questions.
- Improved Technical/Analytical Responses: Focuses on better flow in Turkish technical and analytical answers.
Use Cases
This model is particularly well-suited for applications requiring robust understanding and generation of Turkish text, especially in domains where encyclopedic knowledge and structured information are crucial. It can be beneficial for tasks involving complex Turkish queries and technical documentation.
Limitations
Users should be aware that the model may carry biases from its training data distribution. For critical applications such as legal, health, or financial domains, human expert review is strongly recommended. The model is released under the Apache-2.0 license.