wjbmattingly/comma-qwen-3.5-2b-full
wjbmattingly/comma-qwen-3.5-2b-full is a 2.3 billion parameter Qwen3.5-2B model, fine-tuned by wjbmattingly, specifically designed for transcribing medieval Latin manuscript pages. It produces CATMuS-compliant, line-by-line graphemic transcriptions from page images, focusing on the exact sequence of written letters without editorial intervention. This model excels at accurately converting historical manuscript images into standardized text, achieving a significantly lower Character Error Rate (CER) compared to its base model.
Loading preview...
Overview
wjbmattingly/comma-qwen-3.5-2b-full is a specialized 2.3 billion parameter model, derived from Qwen/Qwen3.5-2B, engineered for the precise transcription of medieval Latin manuscripts. It performs a full fine-tune, meaning it's a standalone checkpoint without adapter or peft dependencies for inference. The model's core function is to generate CATMuS-compliant, line-by-line graphemic transcriptions from page images, preserving the exact written form in modern Latin alphabet without normalization or translation.
Key Capabilities & Performance
- Highly Accurate Transcription: Achieves a Character Error Rate (CER) of 0.1251 (NFD) and a Word Error Rate (WER) of 0.3616 (NFD) on held-out pages, significantly outperforming the base Qwen3.5-2B model (CER 0.8847, WER 1.7516).
- Line-by-Line Output: Provides one line of text for each physical written line in reading order, with a line recall of 0.1647 (NFD).
- Visual Token Handling: Optimized for 2048 visual tokens per page, ensuring consistent resolution with its training data.
- Robustness: Eliminates degenerate and truncated pages during evaluation, indicating improved stability.
Training Details
- Data: Trained on 9,905 pages from the
comma-project/deep-jsonldataset. - Parameters: All 2,213,241,664 parameters were fine-tuned over 3 epochs.
- Hardware: Training was conducted on NVIDIA RTX PRO 6000 Blackwell Server Edition.
Limitations
- CER Interpretation: Differences smaller than ~0.02 CER should be treated as ties due to bf16 training nondeterminism and greedy decoding amplification.
- Batch Invariance: Greedy decoding is not batch-size invariant; comparisons across batch-size changes are invalid.
- Transcription Convention: Remaining errors are primarily due to transcription conventions (e.g., word-division spaces, allographs) rather than genuine letter confusion.
- Specific to Latin & CATMuS: Optimized for Latin-dominant manuscripts under the CATMuS graphemic standard, including the full MUFI superscript-letter repertoire.