ThaiLLM/ThaiLLM-8B-MedApp
ThaiLLM/ThaiLLM-8B-MedApp v2.0.0 is an 8-billion parameter Thai-English medical assistant model developed by ThaiLLM, featuring structured medical tool-calling support. This version is a normalized linear full-weight merge of the original MedApp and ToolUse models, specifically designed to improve tool routing and multi-turn conversational stability. It excels in medical response scoring, citation scoring, and tool selection, making it suitable for research and carefully monitored applications requiring medical conversation and tool integration.
Loading preview...
ThaiLLM-8B-MedApp v2.0.0 Overview
ThaiLLM-8B-MedApp v2.0.0 is an 8-billion parameter Thai-English medical assistant model with enhanced structured medical tool-calling capabilities. Developed by ThaiLLM, this version is a 70/30 normalized linear full-weight merge of the original MedApp and ToolUse models, utilizing the MedApp tokenizer and chat template. The merge was performed using MergeKit 0.1.4 in BF16, with no post-merge fine-tuning.
Key Enhancements and Capabilities
- Improved Tool Routing: The v2.0.0 merge significantly strengthens tool routing behavior, addressing weaknesses found in the original MedApp model for several tool classes.
- Enhanced Conversational Stability: Controlled tests demonstrated improved multi-turn stability, reducing long or repetitive responses during extended conversations.
- Superior Medical Response Quality: Evaluation showed improvements in medical response scoring, citation scoring, and tool selection compared to v1.
- Specific Tool Support: Intended for applications needing Thai medical conversation and routing to tools such as
create_appointment,search_medical_facts,prescreen, andget_health_emergency_contact.
Performance Metrics (v2.0.0 vs. v1)
- med-IQ Evaluation: Achieved 100.00% format correctness, 67.82% citation accuracy, and 75.83% response accuracy, outperforming v1 across all metrics.
- ToolUse Evaluation: Demonstrated near-perfect performance with 99.92% Pass@1 accuracy, 100.00% Trigger F1, and 99.39% Macro F1, significantly improving over v1's 90.36% Pass@1 accuracy.
- Multi-turn Stability: Achieved a 0.00% flag rate in stabilized diagnostic tests, indicating robust conversational flow and minimal output degeneration.
Intended Use Cases
This model is designed for research and carefully monitored applications that require a Thai-English medical assistant capable of engaging in medical conversations and routing to specific medical tools. Users should be aware of limitations regarding medical accuracy and safety, and ensure proper validation and authorization for tool calls.