eth-nlped/Eduardo-27b
Eduardo-27B is a 27 billion parameter math tutor model developed by eth-nlped, fine-tuned from Qwen/Qwen3.8-27B using multi-turn reinforcement learning. Its primary differentiator is its ability to teach students how to solve math problems through Socratic dialogue, rather than providing direct solutions. This model excels at interactive, step-by-step math tutoring and scaffolded dialog systems, making it suitable for research into pedagogical alignment of LLMs.
Loading preview...
Overview
Eduardo-27B is a 27 billion parameter math tutor model, fine-tuned from Qwen/Qwen3.8-27B by eth-nlped using multi-turn reinforcement learning (DPPO). Its core purpose is to teach students how to solve math problems through interactive, Socratic dialogue, rather than simply providing answers. The model was trained in a simulated classroom environment, engaging in multi-turn conversations (up to 22 turns) with a frozen Llama-3.1-8B-Instruct student on Big-Math problems. The training reward mechanism focuses on the student's normalized learning gain on a masked near-transfer post-test, ensuring the tutor guides understanding without directly handing over solutions.
Key Capabilities
- Socratic Math Tutoring: Emphasizes guiding questions and prompting for justification, significantly reducing direct explanations and answer provision compared to its base model.
- High Pedagogical Alignment: Achieves the highest pedagogical reward model (Ped-RM) score and win rate against human teachers on MathDial, outperforming frontier models like Gemini 3.1 Pro.
- Efficiency: Uses 2.4โ6.2 times fewer thinking tokens per turn than frontier models on TutorMoments, indicating efficient reasoning.
- Avoids Over-help: Demonstrates strong performance in avoiding over-help, ensuring students do the cognitive work.
Good For
- Interactive, Step-by-Step Math Tutoring: Ideal for applications requiring an AI to guide users through mathematical problem-solving.
- Socratic and Scaffolded Dialog Systems: Suitable for developing conversational agents that encourage critical thinking and self-discovery.
- Research on Pedagogical LLM Alignment: A valuable research artifact for studying how LLMs can be trained to be effective educators.