Kentucky-Open-Science/KOS-V5-Base
KOS-V5-Base, codenamed "Catbird," is a 3.72 billion parameter medical language model developed by the University of Kentucky and University of Louisville. Trained from scratch on 235.2 billion tokens from a 54-source medical/biomedical corpus, it features a 32,768-token context window and a custom 32k byte-level BPE tokenizer. This base model is designed as a research artifact for medical/biomedical NLP, serving as an initialization for downstream SFT/RL lines rather than an instruction-following or chat model.
Loading preview...
KOS-V5-Base: A From-Scratch Medical Language Model
KOS-V5-Base, or "Catbird," is a 3.72 billion parameter medical language model developed by the University of Kentucky and University of Louisville. Unlike many other models, it was trained entirely from scratch on a specialized 235.2 billion token medical/biomedical corpus, rather than being distilled or continued-pretrained from a general base model. This makes it a unique research artifact for understanding medical language processing.
Key Characteristics & Performance
- Architecture: 3.72B parameters, 36 layers, 32/8 GQA, SwiGLU, tied embeddings.
- Training: One complete epoch over a 54-source medical/biomedical corpus (235.2B tokens).
- Context Length: Trained at 24,576 tokens, with a
max_position_embeddingsof 32,768 tokens. - Tokenizer: Custom 32k byte-level BPE, designed for medical text.
- Evaluation: Achieved 28 wins (top quartile) and 16 losses (bottom quartile) across 94 ranked metrics against a pool of 16 external models, many trained on significantly more data. Notably, it ranked 1st in Bits-per-byte (BPB) on held-out medical text, indicating strong compression efficiency.
- Medical Reasoning: Shows improved performance on closed-book medical MCQ tasks compared to its predecessor, KOS-V4, indicating better medical reasoning capabilities.
Intended Use Cases
This model is a research base model for medical/biomedical NLP. It is ideal for:
- Initialization for SFT/RL: Serving as a foundation for further instruction-tuning or reinforcement learning in medical domains.
- Interpretability Research: Studying model behavior and representations in a specialized medical context.
- Tokenizer/Corpus Research: Investigating the impact of specialized tokenizers and medical corpora.
- From-Scratch Reference: Providing a benchmark against distilled or continued-pretrained medical models.
Important Considerations
- Base Model: KOS-V5-Base is not instruction-tuned and does not follow instructions or engage in chat. It is designed for text completion.
- Known Issues: The tokenizer has limitations with numeric splitting, which may affect tasks involving doses, lab values, or other numerical medical data. It also exhibits low medical single-token rate and struggles with false-confidence tests (hallucination resistance) as a base model without refusal priors.