Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k is a 0.5 billion parameter Qwen2.5-based language model developed by Rumiii, specifically adapted for biomedical applications. It underwent full-parameter continued pre-training on 92k English biomedical abstracts, followed by supervised fine-tuning on a general instruction dataset. This model excels at medical question answering and clinical education, offering a lightweight solution for medical AI prototyping with a 32768 token context length.
Loading preview...
Rumiii/Qwen2.5-0.5B-Med-Post-Trained-92k Overview
This model is a specialized, instruction-tuned variant of the Qwen2.5-0.5B architecture, developed by Rumiii. It has been adapted for the medical domain through a two-stage training process, making it distinct from general-purpose LLMs.
Key Capabilities & Training
- Domain Adaptation: The model underwent full-parameter Continued Pre-Training (CPT) on the
VietAI/vi_pubmeddataset, comprising 92,000 English biomedical abstracts (approximately 23.6 million tokens). This stage significantly improved its understanding of medical terminology and concepts. - Instruction Following: Following CPT, the model was further enhanced with full-parameter Supervised Fine-Tuning (SFT) on 20,000 samples from the
causal-lm/ultrachatdataset, using the Qwen2.5 ChatML format. This fine-tuning improves its ability to follow instructions and generate coherent responses. - Lightweight & Efficient: With 0.5 billion parameters, it is designed to be a lightweight model, suitable for environments with limited computational resources, while still offering specialized medical knowledge.
- Context Length: It supports a substantial context length of 32768 tokens.
Intended Use Cases
- Medical Question Answering: Optimized for providing informative answers to medical queries.
- Clinical Education: Useful for educational purposes related to clinical knowledge.
- Research & Prototyping: Ideal for research into small biomedical language models and rapid prototyping of medical AI applications.
- Demonstration: Showcases an effective CPT + SFT pipeline for domain adaptation on consumer-grade hardware.
Limitations
- Reasoning Depth: Due to its 0.5B parameter size, its reasoning capabilities are more limited compared to much larger models.
- Clinical Accuracy: Outputs are not guaranteed to be clinically accurate and require verification by medical professionals. It is not intended for clinical decision-making or patient care.
- Single-Turn Focus: Primarily trained on single-turn instruction pairs, which may limit its coherence in complex multi-turn conversations.
- Language: English only.