FreedomIntelligence/HuatuoGPT-3-27B
HuatuoGPT-3-27B is a 27 billion parameter medical Large Language Model developed by FreedomIntelligence, built upon the Qwen3.8-27B backbone. It specializes in medical reasoning, adapted using One-stage Policy Optimization (OnePO) which directly optimizes the model for medical tasks without prior supervised fine-tuning. This model is designed to provide medical insights and answers, requiring a specific 'thinking mode' during inference for optimal performance.
Loading preview...
HuatuoGPT-3-27B: A Medical LLM with One-stage Policy Optimization
HuatuoGPT-3-27B is a 27 billion parameter medical Large Language Model (LLM) developed by FreedomIntelligence. It is built on the Qwen3.8-27B architecture and is specifically designed for medical reasoning.
Key Differentiators & Capabilities
- One-stage Policy Optimization (OnePO): Unlike traditional methods that involve supervised fine-tuning, HuatuoGPT-3-27B is adapted to the medical domain using OnePO. This reinforcement learning approach directly optimizes the model for medical tasks, guided by teacher responses that are gradually phased out as the model improves.
- Medical Specialization: The model is explicitly purposed for medical reasoning, making it suitable for applications requiring domain-specific knowledge in healthcare.
- Thinking Mode: For optimal inference, HuatuoGPT-3-27B requires an
enable_thinking=Truesetting. This allows the model to generate internal reasoning before producing its final answer, indicated by</think>. - Resource Availability: FreedomIntelligence has released the training code, a medical RL dataset, and an 8B rubric grader to support its use and development.
When to Use This Model
This model is ideal for applications requiring advanced medical reasoning capabilities. Its unique OnePO training methodology makes it a distinct choice for medical AI research and deployment, particularly where domain-specific adaptation without extensive supervised fine-tuning is desired. It can be deployed using frameworks like vLLM or SGLang.