lastmass/MedGemma-GRPO
MedGemma-GRPO by lastmass is a 4.3 billion parameter fine-tuned version of the MEDGemma model, optimized for clinical case reasoning. It combines Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO) to guide the model in following logical diagnostic paths. This model excels at processing medical case studies, particularly those involving complex cardiovascular, thrombotic, and autoimmune diseases, and is intended for research into enhancing reasoning capabilities in specialized domains.
Loading preview...
MedGemma-GRPO: Enhanced Clinical Reasoning Model
MedGemma-GRPO is a 4.3 billion parameter model developed by lastmass, fine-tuned from the MEDGemma base model. Its core innovation lies in combining Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO) to improve its ability to follow a logical diagnostic path in clinical reasoning tasks. The model was trained on a synthetic dataset of medical case studies, with a notable bias towards complex cardiovascular, thrombotic, and autoimmune diseases.
Key Capabilities
- Logical Diagnostic Reasoning: Utilizes GRPO with a custom rule-based reward function to guide the model towards accurate diagnoses and structured reasoning, while penalizing overly verbose outputs.
- Specialized Medical Domain: Adapted for tasks involving clinical case reasoning, demonstrating proficiency in complex medical scenarios.
- Bias Awareness: Training data has a known bias towards specific disease categories, which is reflected in its performance.
Intended Use & Limitations
This model is strictly for research purposes only, exploring how to enhance reasoning in smaller language models within a specialized domain. It is not a medical device and should never be used for actual medical diagnosis, treatment, or advice due to potential inaccuracies, hallucinations, and biases from its synthetic training data. Always consult a qualified healthcare professional for medical concerns.