buddhist-nlp/qwen35-mitra-dictionary-annotation
The buddhist-nlp/qwen35-mitra-dictionary-annotation is a 9 billion parameter Qwen3.5-based word aligner developed by buddhist-nlp. It is specifically fine-tuned for Sanskrit to English, Tibetan, and Chinese dictionary annotation, identifying and tagging corresponding word spans in translations. This model excels at accurately aligning Sanskrit words with their equivalents in target languages, supporting the creation of attested multilingual dictionaries.
Loading preview...
Overview
The buddhist-nlp/qwen35-mitra-dictionary-annotation is a specialized 9 billion parameter Qwen3.5-based model designed for word alignment in the context of Buddhist texts. Its primary function is to align Sanskrit words with their corresponding translations in English, Tibetan, and Chinese, forming the backbone of the MITRA-dict family of attested dictionaries.
Key Capabilities
- Sanskrit Word Alignment: Given a Sanskrit sentence (segmented and lemmatized) and its translation, the model identifies and tags the translated span for each Sanskrit word.
- Multilingual Support: Supports alignment from Sanskrit to Tibetan, Chinese, and English. It also handles Chinese to Tibetan and English to Tibetan directions.
- Dictionary Annotation: Facilitates the creation of high-quality, attested Sanskrit dictionaries by programmatically linking source and target words.
- Prompt-based Tagging: Utilizes a specific prompt format to insert numbered tags into the target translation, marking the word(s) that render each numbered source word.
Training and Performance
The model underwent full fine-tuning of MITRA Qwen3.5-9B, which itself was continuously pre-trained on Sanskrit, Tibetan, Buddhist Chinese, and Pāli. It was fine-tuned for three epochs on 9,366 sentence pairs labeled by a commercial LLM. Evaluation against 1,489 held-out sentence pairs shows strong agreement with the teacher model:
- Sanskrit to Tibetan: 88.9% F1
- Sanskrit to Chinese: 89.2% F1
- Sanskrit to English: 86.9% F1
Usage Notes
The model uses the Qwen3_5ForConditionalGeneration layout and is intended for greedy decoding, stopping at the first newline. The vision tower present in the base Qwen3.5 architecture is unused in this specialized application. Further details on the data and evaluation protocols can be found on the Dharmamitra Lexicon GitHub.