AITC-URAx/AITC-URAx

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 2, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AITC-URAx/AITC-URAx is an 8 billion parameter Vietnamese instruction-tuned LLM developed by AITC-URAx, specifically optimized for the Vietnamese K-12 education environment. It excels in reasoning, pedagogical interaction, and ethical responses, utilizing a unique three-stage SFT-GRPO-DPO pipeline with consistent Vietnamese Chain-of-Thought (CoT) training. This model is designed to support 12 educational tasks, including QA reasoning, personalized learning, and psychological support, with a focus on reducing hallucination.

Loading preview...

AITC-URAx: An LLM for Vietnamese K-12 Education

AITC-URAx is an 8 billion parameter instruction-tuned large language model (LLM) developed by AITC-URAx, specifically engineered for the Vietnamese K-12 educational context. It is built upon four core pillars: Knowledge (TÀI) covering textbook content and multi-step reasoning; Ethics (ĐỨC) ensuring safety, responsibility, and appropriate refusal in sensitive situations; Pedagogy (SƯ PHẠM) adopting a natural teacher-student tone with clear, evocative explanations; and Reasoning through a consistent Vietnamese Chain-of-Thought (CoT) structure ("Bước 1 – Bước 2 – Kết luận") applied across all training stages.

Key Capabilities

  • Comprehensive Task Support: Handles 12 educational tasks, including QA reasoning, error correction, idea generation, personalized learning, psychological support, question generation, automated grading, and teaching material creation.
  • Advanced Reasoning: Deeply invested in QA reasoning, multi-turn interactive learning, and school-based psychological/safety support.
  • Pedagogical Tools: Excels at creating exams, MCQs, and graded assignments with explanations.
  • Reduced Hallucination: Significantly lowers hallucination rates through Unanswerable QA data and unique DPO training with correct-vs-incorrect reasoning pairs.

Unique Training Strategy

The model's primary differentiator is its CoT training applied xuyên suốt (throughout) all three stages: Supervised Fine-Tuning (SFT), Guided Reinforcement Learning with Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Unlike models that only use CoT in SFT, AITC-URAx integrates CoT into GRPO (rewarding good reasoning chains) and DPO (learning to distinguish correct vs. incorrect reasoning), enhancing reliability and reducing factual errors.

Good for

  • Vietnamese educational applications: Tailored for K-12 curriculum and pedagogical interactions.
  • Reasoning-intensive tasks: Benefits from consistent CoT training.
  • Ethical and safe AI interactions: Designed to handle sensitive topics responsibly within an educational setting.
  • Automated content generation for teachers: Creates teaching materials, quizzes, and provides grading with explanations.