MooreMuaMu/qwen35-27b-ancient-rl-r32-step200

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MooreMuaMu/qwen35-27b-ancient-rl-r32-step200 is a 27 billion parameter Qwen3.5-based language model, fine-tuned using Reinforcement Learning (RL) with a rank-32 GRPO checkpoint. This model specializes in processing and generating content related to ancient languages, specifically Tibetan, Traditional Mongolian, and Uyghur. It is designed to improve answer extraction and semantic understanding in these specialized linguistic contexts, offering enhanced performance for tasks requiring precise information retrieval from ancient texts.

Loading preview...

Model Overview

MooreMuaMu/qwen35-27b-ancient-rl-r32-step200 is a 27 billion parameter language model built upon the Qwen3.5 architecture. It has undergone Reinforcement Learning (RL) fine-tuning using a rank-32 GRPO checkpoint at step 200, with the LoRA adapter merged into the base model. This full merged bfloat16 model is provided with safetensors shards and necessary tokenizer/processor/chat-template files for direct use with the transformers library.

Key Capabilities and Specialization

This model is specifically trained and optimized for tasks involving ancient languages, including Tibetan, Traditional Mongolian, and Uyghur. Evaluation results indicate several key improvements:

  • Enhanced Answer Presence: The model significantly improves the likelihood of extracting an answer, with a +4.67% delta in "Has answer" metric and a +5.0 pp increase in "Nonempty extracted answer" compared to the base model.
  • Semantic Understanding: It shows positive movement in overall semantic metrics, including a +0.0274 increase in BERTScore-F1 and a +0.0078 increase in Semantic composite score.
  • Ancient Language Processing: While performance varies across specific tasks and languages, it demonstrates notable gains in Corpus BLEU-2 for Tibetan translation (+7.06 pp), Traditional Mongolian annotation (+5.94 pp), and Uyghur annotation (+7.72 pp).

Training Details

The model was fine-tuned from a Qwen3.5-27B ancient-language stage-2 base model. The LoRA checkpoint utilized a rank/alpha of 32/32, and the merge was performed in bfloat16 dtype. The serialization uses safetensors with a maximum shard size of 5GB.

Intended Use Cases

This model is particularly well-suited for research and applications requiring specialized processing of ancient texts in Tibetan, Traditional Mongolian, and Uyghur. Its strengths lie in tasks such as:

  • Information Extraction: Accurately identifying and extracting answers from complex ancient language documents.
  • Semantic Analysis: Understanding the meaning and context within these specialized linguistic domains.
  • Translation and Annotation Support: Assisting with translation and annotation tasks for the specified ancient languages, though performance can be mixed depending on the specific language and task.