AITC-URAx/AITC-URAx
AITC-URAx/AITC-URAx is an 8 billion parameter Vietnamese instruction-tuned LLM developed by AITC-URAx, specifically optimized for the Vietnamese K-12 education environment. It excels in reasoning, pedagogical interaction, and ethical responses, utilizing a unique three-stage SFT-GRPO-DPO pipeline with consistent Vietnamese Chain-of-Thought (CoT) training. This model is designed to support 12 educational tasks, including QA reasoning, personalized learning, and psychological support, with a focus on reducing hallucination.
Loading preview...
AITC-URAx: An LLM for Vietnamese K-12 Education
AITC-URAx is an 8 billion parameter instruction-tuned large language model (LLM) developed by AITC-URAx, specifically engineered for the Vietnamese K-12 educational context. It is built upon four core pillars: Knowledge (TÀI) covering textbook content and multi-step reasoning; Ethics (ĐỨC) ensuring safety, responsibility, and appropriate refusal in sensitive situations; Pedagogy (SƯ PHẠM) adopting a natural teacher-student tone with clear, evocative explanations; and Reasoning through a consistent Vietnamese Chain-of-Thought (CoT) structure ("Bước 1 – Bước 2 – Kết luận") applied across all training stages.
Key Capabilities
- Comprehensive Task Support: Handles 12 educational tasks, including QA reasoning, error correction, idea generation, personalized learning, psychological support, question generation, automated grading, and teaching material creation.
- Advanced Reasoning: Deeply invested in QA reasoning, multi-turn interactive learning, and school-based psychological/safety support.
- Pedagogical Tools: Excels at creating exams, MCQs, and graded assignments with explanations.
- Reduced Hallucination: Significantly lowers hallucination rates through Unanswerable QA data and unique DPO training with correct-vs-incorrect reasoning pairs.
Unique Training Strategy
The model's primary differentiator is its CoT training applied xuyên suốt (throughout) all three stages: Supervised Fine-Tuning (SFT), Guided Reinforcement Learning with Policy Optimization (GRPO), and Direct Preference Optimization (DPO). Unlike models that only use CoT in SFT, AITC-URAx integrates CoT into GRPO (rewarding good reasoning chains) and DPO (learning to distinguish correct vs. incorrect reasoning), enhancing reliability and reducing factual errors.
Good for
- Vietnamese educational applications: Tailored for K-12 curriculum and pedagogical interactions.
- Reasoning-intensive tasks: Benefits from consistent CoT training.
- Ethical and safe AI interactions: Designed to handle sensitive topics responsibly within an educational setting.
- Automated content generation for teachers: Creates teaching materials, quizzes, and provides grading with explanations.