zero9tech/Qwen3-8B-Data-Science-Insight-TR-7.6K

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The zero9tech/Qwen3-8B-Data-Science-Insight-TR-7.6K is an 8 billion parameter Qwen3-based language model developed by Zero9 Tech, fine-tuned for data mining and applied data science decision support. It features a 32K context length and is specifically adapted for Turkish language reasoning through continued pre-training on Wikipedia. This model excels at generating decision-oriented responses, including method selection, alternative comparisons, risk signaling, and validation steps.

Loading preview...

Model Overview

The zero9tech/Qwen3-8B-Data-Science-Insight-TR-7.6K is an 8 billion parameter model developed by Zero9 Tech, specifically engineered for data mining and applied data science decision support. It leverages the Qwen3 architecture and has undergone a specialized training regimen to optimize its performance in these domains.

Key Capabilities

  • Turkish Language Adaptation: The model underwent continued pre-training (CPT) with approximately 10% of the wikimedia/wikipedia dataset to adapt its reasoning capabilities to Turkish.
  • Domain Expertise: It was further fine-tuned using the murataksit34/veri-bilimci-diyalog-8k-tr dataset, focusing on expert dialogues in data science.
  • Decision-Oriented Responses: Optimized to produce actionable insights and support decision-making, including:
    • Method selection
    • Comparison of alternatives
    • Identification of risk signals
    • Validation steps

Training Details

The model's training involved an initial Turkish adaptation phase followed by a specialized Supervised Fine-Tuning (SFT) phase. The SFT dataset, murataksit34/veri-bilimci-diyalog-8k-tr, comprises 7,656 records, split into 6,124 for training and 1,532 for testing, with high uniqueness ratios for assistant responses.

Good For

This model is particularly well-suited for applications requiring nuanced, decision-focused outputs in Turkish within the fields of data mining and applied data science.