SemanticAlignment/Mistral-v0.1-Italian-FVT
Mistral-v0.1-Italian-FVT is a 7 billion parameter, continually trained Mistral model adapted for the Italian language by SapienzaNLP, ISTI-CNR, and ILC-CNR. It utilizes an optimized transformer architecture and a substituted tokenizer, similar to Minerva-3B, to enhance efficiency for Italian text. This model is specifically fine-tuned on a dataset skewed towards Italian, making it highly effective for generative tasks in Italian.
Loading preview...
Model Overview
SemanticAlignment/Mistral-v0.1-Italian-FVT is a 7 billion parameter generative language model, adapted from the Mistral-7B-Base-v0.1 architecture. Developed by SapienzaNLP, ISTI-CNR, and ILC-CNR, this model has undergone continuous training with a substituted tokenizer, aligning it with the tokenizer used by Minerva-3B.
Key Adaptations and Training
- Tokenizer Substitution: The model's tokenizer has been replaced to optimize its performance for the Italian language, enhancing token fertility and efficiency.
- Data Skewed Towards Italian: Training involved a custom dataset derived from CulturaX, with a significant emphasis on Italian content. Specifically, it was trained on 9 billion Italian tokens and 3 billion English tokens from CulturaX.
- Optimized Transformer Architecture: It leverages the efficient Mistral-7B-v0.1 architecture, making it suitable for various text generation tasks.
Use Cases
This model is particularly well-suited for applications requiring high-quality text generation and understanding in Italian, benefiting from its specialized tokenizer and Italian-centric training data. Developers can integrate it using the Hugging Face Transformers library for conversational inference and other generative tasks.