wjbmattingly/comma-qwen-3.5-2b-full

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

wjbmattingly/comma-qwen-3.5-2b-full is a 2.3 billion parameter Qwen3.5-2B model, fine-tuned by wjbmattingly, specifically designed for transcribing medieval Latin manuscript pages. It produces CATMuS-compliant, line-by-line graphemic transcriptions from page images, focusing on the exact sequence of written letters without editorial intervention. This model excels at accurately converting historical manuscript images into standardized text, achieving a significantly lower Character Error Rate (CER) compared to its base model.

Loading preview...

Overview

wjbmattingly/comma-qwen-3.5-2b-full is a specialized 2.3 billion parameter model, derived from Qwen/Qwen3.5-2B, engineered for the precise transcription of medieval Latin manuscripts. It performs a full fine-tune, meaning it's a standalone checkpoint without adapter or peft dependencies for inference. The model's core function is to generate CATMuS-compliant, line-by-line graphemic transcriptions from page images, preserving the exact written form in modern Latin alphabet without normalization or translation.

Key Capabilities & Performance

  • Highly Accurate Transcription: Achieves a Character Error Rate (CER) of 0.1251 (NFD) and a Word Error Rate (WER) of 0.3616 (NFD) on held-out pages, significantly outperforming the base Qwen3.5-2B model (CER 0.8847, WER 1.7516).
  • Line-by-Line Output: Provides one line of text for each physical written line in reading order, with a line recall of 0.1647 (NFD).
  • Visual Token Handling: Optimized for 2048 visual tokens per page, ensuring consistent resolution with its training data.
  • Robustness: Eliminates degenerate and truncated pages during evaluation, indicating improved stability.

Training Details

  • Data: Trained on 9,905 pages from the comma-project/deep-jsonl dataset.
  • Parameters: All 2,213,241,664 parameters were fine-tuned over 3 epochs.
  • Hardware: Training was conducted on NVIDIA RTX PRO 6000 Blackwell Server Edition.

Limitations

  • CER Interpretation: Differences smaller than ~0.02 CER should be treated as ties due to bf16 training nondeterminism and greedy decoding amplification.
  • Batch Invariance: Greedy decoding is not batch-size invariant; comparisons across batch-size changes are invalid.
  • Transcription Convention: Remaining errors are primarily due to transcription conventions (e.g., word-division spaces, allographs) rather than genuine letter confusion.
  • Specific to Latin & CATMuS: Optimized for Latin-dominant manuscripts under the CATMuS graphemic standard, including the full MUFI superscript-letter repertoire.