surrey-nlp/diallm-qwen-grpo-aus

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The surrey-nlp/diallm-qwen-grpo-aus model is an 8 billion parameter Qwen 3-8B variant, fine-tuned for Australian English (en-AU) dialect adaptation. It was continually pretrained on the International Corpus of English and further adapted using dialect-specific Supervised Fine-Tuning (SFT) on Multi-VALUE-transformed en-AU preference data. The model employs Generative Reinforcement Learning from Policy Optimization (GRPO) with target-variety preference pairs to enhance its robustness and generation capabilities for Australian English.

Loading preview...

DiaLLM - Qwen 3-8B - Australian English Adaptation

This model, surrey-nlp/diallm-qwen-grpo-aus, is an 8 billion parameter variant of the Qwen 3-8B base model, specifically adapted for Australian English (en-AU). It is a key component of the "DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation" research.

Key Adaptation Details

  • Target Variety: Australian English (en-AU).
  • Adaptation Method: The model underwent an "explicit" (variety-targeted) adaptation thread.
  • Training Process:
    • Initially continually pretrained on the International Corpus of English (18 varieties, ~20M tokens).
    • Further adapted via dialect-specific Supervised Fine-Tuning (SFT) using Multi-VALUE-transformed en-AU preference data.
    • Aligned using Generative Reinforcement Learning from Policy Optimization (GRPO) with target-variety preference pairs.
  • Base Model: It is a fine-tuned version of jordanpainter/diallm-qwen-sft-aus.

Use Cases

This model is particularly suited for applications requiring language generation or understanding with a strong emphasis on Australian English linguistic nuances and expressions. It aims to bridge the robustness-generation gap in dialect adaptation, making it valuable for research and development in localized NLP tasks.

Further Resources