srirag/alexmt-gemma4-e4b-dpo-bwspbleu-lr2e6

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:cc-by-nc-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

srirag/alexmt-gemma4-e4b-dpo-bwspbleu-lr2e6 is a 7.9 billion parameter Gemma-4-E4B-it based model fine-tuned by srirag for context-aware English to dialectal Arabic and dialectal Arabic to English translation. It specializes in conversational turns across nine Arabic varieties, leveraging Direct Preference Optimization (DPO) on the Alexandria dataset. This model demonstrates improved spBLEU and chrF++ scores for both translation directions compared to its SFT base, making it suitable for nuanced machine translation in conversational AI.

Loading preview...

Overview

This model, srirag/alexmt-gemma4-e4b-dpo-bwspbleu-lr2e6, is a 7.9 billion parameter language model built upon the Gemma-4-E4B-it architecture. It has been specifically fine-tuned using Direct Preference Optimization (DPO) to excel in context-aware English ↔ dialectal Arabic translation for conversational turns.

Key Capabilities

  • Dialectal Arabic Translation: Supports translation between English and nine dialectal Arabic varieties (Egyptian, Jordanian, Lebanese, Moroccan, Omani, Palestinian, Saudi, Syrian, Yemeni).
  • Context-Aware: Utilizes up to three prior turns as context, along with domain, participants, and gender direction, to inform translations.
  • DPO Fine-tuning: Enhanced through DPO (beta 0.1, lr 2e-6, 3 epochs) on preference pairs derived from the Alexandria dataset, leading to improved translation quality.
  • Performance: Achieves notable spBLEU and chrF++ scores on the Alexandria public test split, outperforming its SFT base model. For instance, it scores 30.12 spBLEU for en→dialect and 53.14 spBLEU for dialect→en.

Usage and Limitations

The model expects Alexandria's JSON prompt format for input and outputs {"translation": "..."}. It's important to note that the DPO learning rate was compared on the test split, and results are not directly comparable to the Alexandria paper's tables due to different setups. Only overlap metrics (spBLEU, chrF++) have been evaluated; dialect fidelity has not been assessed.