Yigit-Karaman/Cozum-4B
Cozum-4B is a 4.5 billion parameter language model developed by Yigit-Karaman, built on the Qwen architecture. It is specifically optimized for Turkish cultural context, instruction following, and mathematical problem-solving. The model excels at providing culturally aware and accurate responses for native Turkish speakers, going beyond simple translation. It is fine-tuned with datasets like Aya Turkish Filtered and a Turkish translation of GSM8K to enhance its reasoning and localized understanding.
Loading preview...
Cozum-4B: Turkish-Optimized Language Model
Cozum-4B is a 4.5 billion parameter language model developed by Yigit-Karaman, based on the Qwen architecture. It is uniquely optimized for the Turkish cultural context, Turkish instruction following, and mathematical problem-solving. The model aims to provide natural, culturally aware, and highly accurate responses for native Turkish speakers, moving beyond basic translation.
Key Capabilities & Training
- Cultural Nuance: Fine-tuned to understand and generate text with local Turkish nuances, idioms, and cultural knowledge.
- Instruction Following: Enhanced ability to follow instructions in Turkish, ensuring alignment with user expectations.
- Mathematical Reasoning: Improved step-by-step mathematical problem-solving capabilities, trained on a Turkish translation of the GSM8K dataset.
- Dataset Curation: Underwent Supervised Fine-Tuning (SFT) using a curated mix of localized datasets, including a filtered subset of the Aya Turkish Filtered dataset, the broader Aya Dataset, and Alpaca Turkish Combined.
- Architecture: Utilizes the Qwen architecture and is trained using LLaMA-Factory.
When to Use This Model
Cozum-4B is particularly well-suited for applications requiring:
- Culturally sensitive AI interactions in Turkish.
- Accurate instruction following for Turkish users.
- Mathematical problem-solving in a Turkish context.
- Generating natural and contextually appropriate Turkish text.