surfiniaburger/medgemma-4b-dipg-safety-v2
surfiniaburger/medgemma-4b-dipg-safety-v2 is a 1 billion parameter MedGemma 4B model fine-tuned by surfiniaburger using Group Relative Policy Optimization (GRPO) on the DIPG Safety Gym benchmark. With a 32768 token context length, this model is specifically optimized for medical safety, focusing on reducing hallucinations and generating evidence-grounded responses in medical contexts. It is designed to enhance the reliability of AI in sensitive medical applications.
Loading preview...
Model Overview
This model, surfiniaburger/medgemma-4b-dipg-safety-v2, is a specialized version of the 1 billion parameter MedGemma 4B instruction-tuned model. It has been fine-tuned by surfiniaburger to enhance safety and reliability, particularly within medical domains.
Key Capabilities
- Medical Safety Focus: Specifically trained to improve safety in medical applications.
- Hallucination Reduction: Engineered to minimize the generation of incorrect or fabricated information.
- Evidence-Grounded Responses: Aims to produce responses that are supported by factual evidence.
- GRPO Fine-tuning: Utilizes Group Relative Policy Optimization (GRPO) with LoRA (rank=64, alpha=64) for targeted safety improvements.
- DIPG Safety Gym Integration: Part of the broader DIPG Safety Gym ecosystem, indicating its role in a dedicated medical safety benchmark.
When to Use This Model
This model is particularly well-suited for use cases requiring high reliability and safety in medical contexts, where the reduction of hallucinations and the generation of evidence-based information are critical. Its specialized training makes it a strong candidate for applications involving sensitive medical data or decision support.