Nalandadata/nalanda-qwen-7b-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Gated Featherless Exclusive Cold

Nalandadata/nalanda-qwen-7b-grpo is a fine-tuned Qwen 2.5 7B Instruct model developed by Nalanda Data, specialized for solving Indian competitive exam questions (JEE Mains, JEE Advanced, NEET UG) across Physics, Chemistry, Mathematics, and Biology. It utilizes a two-stage training pipeline, including Group Relative Policy Optimization (GRPO), to achieve significant accuracy improvements on these exams while preserving general reasoning capabilities. The model demonstrates an overall 9.1 percentage point accuracy improvement on held-out JEE/NEET MCQs compared to its baseline.

Loading preview...

Nalanda Qwen 2.5 7B GRPO: Specialized for Indian Competitive Exams

This model, developed by Nalanda Data, is a fine-tuned version of Qwen 2.5 7B Instruct, specifically optimized for answering questions from Indian competitive exams like JEE Mains, JEE Advanced, and NEET UG. It covers Physics, Chemistry, Mathematics, and Biology.

Key Capabilities & Training

The model was trained using an innovative two-stage pipeline:

  • Stage 1: Light Supervised Fine-Tuning (SFT): Briefly introduced domain vocabulary and question formats using a mix of JEE/NEET questions and general instruction data (SlimOrca).
  • Stage 2: Group Relative Policy Optimization (GRPO): This crucial stage, inspired by recent research, trained the model to arrive at correct answers through its own reasoning. Unlike standard SFT, GRPO prevents catastrophic forgetting by rewarding correctness, format compliance, and reasoning quality, rather than forcing specific solution patterns.

Performance & Preservation

Nalanda Qwen 2.5 7B GRPO shows substantial improvements on held-out JEE/NEET exam questions:

  • Overall Accuracy: Improved by +9.1 percentage points (pp) to 69.6% from a 60.5% baseline.
  • Subject-specific gains: Physics (+14.0pp), Chemistry (+10.0pp), Mathematics (+8.5pp), Biology (+4.0pp).
  • Benchmark Preservation: Crucially, the model maintains or slightly improves performance on public benchmarks like GSM8K, ARC-Challenge, and MMLU-Physics/Chemistry, indicating that its general reasoning abilities are fully preserved and even enhanced.

Use Cases

This model is ideal for applications requiring high accuracy on Indian competitive exam questions, particularly those involving multiple-choice questions with detailed reasoning. It leverages a dataset of 116,831 expert-curated JEE/NEET exam questions from Nalanda Data, including LaTeX mathematical notation. The model is released under the Apache 2.0 license.