ATH-MaaS/Marco-LLM-AR-V2

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 10, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Marco-LLM-AR-V2 is a 7.6 billion parameter language model from ATH-MaaS, specifically fine-tuned for common languages and dialects used in the Arab world. It is based on the Transformer architecture with SwiGLU activation and an improved tokenizer adaptive to multiple Arabic dialects. The model underwent extensive continued pretraining on approximately 50 billion tokens, enhancing its capabilities in targeted languages while maintaining general benchmark competitiveness.

Loading preview...

Marco-LLM-AR-V2: Enhanced Arabic Language Model

Marco-LLM-AR-V2 is a 7.6 billion parameter model from the Marco-LLM-AR series, specifically designed and fine-tuned for languages prevalent in the Arab world, including Modern Standard Arabic and various dialects. This model has undergone extensive continued pretraining on a substantial dataset of approximately 50 billion tokens, which significantly enhances its performance in these targeted languages while ensuring competitive general language capabilities.

Key Capabilities & Architecture

  • Arabic Language Specialization: Optimized for Modern Standard Arabic and multiple Arabic dialects through dedicated pretraining.
  • Transformer Architecture: Utilizes a Transformer architecture with SwiGLU activation, attention QKV bias, and group query attention.
  • Adaptive Tokenizer: Features an improved tokenizer specifically adapted to handle the nuances of multiple Arabic dialects and forms.

Usage Recommendations

This base model is not intended for direct text generation without further adaptation. It is recommended to apply post-training methods such as Supervised Fine-tuning (SFT), Reinforcement Learning with Human Feedback (RLHF), or additional continued pretraining to tailor the model for specific downstream applications and use cases. For more details, refer to the Hugging Face page and the associated research paper: Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement.