Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k
Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k is a 0.5 billion parameter Qwen2.5-based language model, developed by Rumiii, specifically adapted for biomedical applications. It was created through full-parameter continued pre-training on 92k English PubMed abstracts, followed by supervised fine-tuning on a general instruction dataset. This model excels at medical question answering and clinical education, offering a lightweight solution for research and prototyping in the biomedical domain.
Loading preview...
Model Overview
This model, Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k, is a 0.5 billion parameter variant of Qwen/Qwen2.5-0.5B, specifically adapted for medical and biomedical text. It was developed by Rumiii through a two-stage training process. The first stage involved full-parameter continued pre-training (CPT) on approximately 23.6 million tokens from the VietAI/vi_pubmed dataset, comprising 92,000 English abstracts. This was followed by a second stage of full-parameter supervised fine-tuning (SFT) on 20,000 samples from the causal-lm/ultrachat dataset, utilizing the Qwen2.5 ChatML template.
Key Capabilities
- Domain Adaptation: Specialized for biomedical text through continued pre-training on PubMed abstracts.
- Instruction Following: Fine-tuned to respond to general instructions, particularly effective with a recommended system prompt for medical AI assistance.
- Lightweight: At 0.5 billion parameters, it is suitable for research and prototyping on consumer hardware.
- Medical Question Answering: Designed to provide clear, direct, and informative answers to medical questions.
Intended Use Cases
- Medical Question Answering and Clinical Education: Ideal for generating factual medical information.
- Biomedical Research: Useful for exploring small language models in the biomedical field.
- Lightweight AI Prototyping: Suitable for developing and demonstrating medical AI applications.
Limitations
Due to its 0.5 billion parameter size, the model has limited reasoning depth and basic multi-turn coherence. Clinical accuracy is not guaranteed, and all outputs require expert verification. It is English-only and may produce inconsistent responses to simple greetings.