lukasdrews/Gemma-4-E2B-IT-SFT-RLVR-Medical
The lukasdrews/Gemma-4-E2B-IT-SFT-RLVR-Medical model is a 5.1 billion parameter Gemma-4-E2B-it variant, fine-tuned by lukasdrews using Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLVR) specifically on the PubMedQA dataset. This model is optimized for medical question answering, demonstrating enhanced performance in in-domain medical contexts and zero-shot transfer capabilities to related medical tasks. With a 32768 token context length, it is designed for specialized applications requiring accurate medical information retrieval and reasoning.
Loading preview...
Model Overview
The lukasdrews/Gemma-4-E2B-IT-SFT-RLVR-Medical is a 5.1 billion parameter language model based on the Gemma-4-E2B-it architecture. It has been specifically fine-tuned by lukasdrews using a combination of Supervised Fine-Tuning (SFT) and Reinforcement Learning with Value Representation (RLVR) on the PubMedQA dataset. This specialized training aims to significantly improve its performance in medical question-answering tasks.
Key Capabilities and Performance
This model demonstrates notable improvements in medical domain understanding:
- Enhanced Medical QA: Achieves 73.10% accuracy on PubMedQA (in-domain), a substantial increase from the base Gemma-4-E2B-it model's 58.10%.
- Zero-Shot Transfer: Shows improved zero-shot transfer capabilities on MedQA-USMLE, reaching 43.05% compared to the base model's 29.54%.
- Quantization Robustness: Performance remains strong even with various quantization levels (e.g., Q8_0 at 72.40% on PubMedQA).
Use Cases
This model is particularly well-suited for:
- Medical Question Answering: Ideal for applications requiring accurate responses to medical queries, leveraging its specialized training on PubMedQA.
- Biomedical Information Retrieval: Can be used in systems that need to process and understand complex medical texts.
- Research and Development: Provides a strong foundation for further fine-tuning or research in medical AI applications.