nassimjp/LFM2.5-1.2B-Instruct-Pashto
The nassimjp/LFM2.5-1.2B-Instruct-Pashto is a 1.2 billion parameter instruction-tuned language model, building upon the LiquidAI/LFM2.5-1.2B-Base architecture. This model is specifically designed and optimized for the Pashto language, incorporating a specialized tokenizer for Pashto atomic units. It is intended for applications requiring high-quality Pashto language understanding and generation, leveraging supervised fine-tuning on Pashto instruction datasets.
Loading preview...
Model Overview
The nassimjp/LFM2.5-1.2B-Instruct-Pashto is a 1.2 billion parameter instruction-tuned language model, derived from the LiquidAI/LFM2.5-1.2B-Base architecture. Its primary distinction lies in its specialized focus on the Pashto language, incorporating a custom tokenizer designed to handle unique Pashto atomic units such as "ښ", "څ", "ځ", and others.
Key Capabilities
- Pashto Language Processing: Optimized for understanding and generating text in Pashto.
- Instruction Following: Fine-tuned with high-quality Pashto instruction datasets to respond to prompts effectively.
- Specialized Tokenization: Utilizes a tokenizer specifically adapted for the nuances of the Pashto script, enhancing linguistic accuracy.
Training Methodology
The model's development involved:
- Continued Pretraining (CPT): Requires extensive, deduped, and domain-balanced Pashto corpora.
- Supervised Fine-Tuning (SFT): Leverages high-quality Pashto instruction datasets formatted for chat applications.
Use Cases
This model is particularly well-suited for applications requiring robust Pashto language capabilities, including:
- Pashto-specific chatbots and conversational AI.
- Text generation and summarization in Pashto.
- Language understanding tasks for Pashto content.
Licensing
The model operates under the same license terms as its base model, LiquidAI/LFM2.5-1.2B-Base. Users should consult the original model card for detailed licensing information.