con-cord/Mod2-with-ref
con-cord/Mod2-with-ref is a 4.3 billion parameter medical LLM-as-a-Judge model based on Gemma-3-4B, fine-tuned for evaluating generated medical responses. This version specifically operates with expert reference answers, making it suitable for assessing medical answer quality against established clinical criteria. Its primary strength lies in its specialized application for medical response evaluation.
Loading preview...
Overview
con-cord/Mod2-with-ref is a specialized medical Large Language Model (LLM) designed to function as an "LLM-as-a-Judge." Built upon the Gemma-3-4B architecture, this model has been fine-tuned for the critical task of evaluating the quality of generated medical answers. With 4.3 billion parameters and a context length of 32768 tokens, it offers a robust solution for automated assessment in medical contexts.
Key Capabilities
- Medical Response Evaluation: Specifically trained to assess the quality of AI-generated medical answers.
- Reference-Based Judging: This particular version (
with-ref) utilizes expert reference answers to guide its evaluation process, ensuring assessments are grounded in established clinical knowledge. - Clinical Criteria Adherence: Designed to evaluate responses according to predefined clinical evaluation criteria.
- Base Model: Leverages the Gemma-3-4B base model, enhanced with Transformers, PEFT/LoRA, and TRL frameworks for its specialized task.
Good For
- Automated Quality Assurance: Ideal for systems requiring automated evaluation of medical text generation.
- Clinical Content Validation: Useful for validating the accuracy and appropriateness of AI-generated medical information against expert-provided references.
- Research in Medical AI: Provides a tool for researchers developing and testing medical LLMs, offering a method to benchmark response quality.