ya-yje-krasni/qwen3-0.6b-russian-dialogues

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The ya-yje-krasni/qwen3-0.6b-russian-dialogues model is a fully fine-tuned Qwen3-0.6B variant, developed by ya-yje-krasni, specifically optimized for generating responses in Russian dialogues. This 0.8 billion parameter model leverages a 32,768 token context length and was trained on the Den4ikAI/russian_dialogues dataset. It excels at producing short, conversational Russian replies, making it suitable for interactive dialogue systems where factual accuracy is not the primary concern.

Loading preview...

Model Overview

The ya-yje-krasni/qwen3-0.6b-russian-dialogues is a specialized language model based on the Qwen3-0.6B architecture. It has undergone a full fine-tuning process, meaning all model parameters were updated, rather than using LoRA or other adapter methods. The model is designed for generating responses within Russian dialogue contexts.

Key Characteristics

  • Base Model: Qwen/Qwen3-0.6B (specifically unsloth/Qwen3-0.6B-Base)
  • Training: Full fine-tuning on all parameters.
  • Dataset: Trained exclusively on the Den4ikAI/russian_dialogues dataset.
  • Architecture: Qwen3ForCausalLM with 28 layers, 1024 hidden size, 16 attention heads, and 8 KV-heads.
  • Context Length: Supports a substantial context window of 32,768 tokens.
  • Vocabulary: Utilizes a Qwen2Tokenizer with a vocabulary size of 151,936 tokens.
  • Language: Optimized specifically for Russian.
  • Prompt Format: Employs a distinct prompt template: <user replica>\n### Ответ: where generation continues after ### Ответ:.

Use Cases and Limitations

This model is best suited for applications requiring short, conversational responses in Russian dialogues. Due to its 0.6 billion parameter size and training focus, it is prone to generating factually incorrect information and is not recommended for tasks demanding factual accuracy.