mikelalda/qwen3.8-27b_euskara
The mikelalda/qwen3.8-27b_euskara model is a 27 billion parameter Qwen3.8-27B base model fine-tuned by mikelalda, specifically optimized for Euskara (Basque) with Spanish and English as support languages. It leverages a hybrid gated delta rule attention mechanism and was trained using LoRA (PEFT) on approximately 74,000 examples from HiTZ corpora. This model excels in instruction-following conversations, multi-directional translation, and text-based tasks like summarization and answerability, making it suitable for applications requiring strong multilingual capabilities, particularly in Basque.
Loading preview...
Model Overview
mikelalda/qwen3.8-27b_euskara is a fine-tuned version of the Qwen3.8-27B model, specifically adapted for Euskara (Basque), with Spanish and English as supporting languages. Developed by mikelalda, this 27 billion parameter model utilizes a hybrid gated delta rule attention mechanism, requiring flash-linear-attention for optimal performance. It was fine-tuned using LoRA (PEFT) with Unsloth and TRL's SFTTrainer.
Key Capabilities
- Multilingual Instruction Following: Trained on conversational instructions in Basque, Spanish, and English.
- Multi-directional Translation: Capable of translating between Basque, Spanish, and English, with specific focus on
eu ↔ esandeu ↔ enpairs. - Text-based Tasks: Excels at summarization and determining answerability based on provided text.
- Context Length: Supports a sequence length of 4096 tokens during training.
Training Details
The model was trained for one epoch on approximately 74,000 examples from HiTZ corpora, including:
HiTZ/magpie-en-eu-reasoning-instructions-qwen3: For conversational instructions.HiTZ/ALIA_syntethic_MT_V2: For synthetic multi-directional translation data.HiTZ/RAG_eu: Used for text-based summarization and answerability tasks. Note: This dataset was also an evaluation bank, leading to potential evaluation contamination if used for benchmarking.
Limitations and Considerations
- Unevaluated Performance: No benchmarks have been run; performance claims are not yet validated.
- Untrained Reasoning: The SFT was performed with
enable_thinking=False, meaning the base model's reasoning mode was not reinforced and may have degraded. - Multimodal Path: While the base Qwen3.8-27B is multimodal, the vision path was not trained or verified in this fine-tuning.
- Domain Bias: Training data from sources like Berria (press), official gazettes, and parliamentary acts results in a formal, institutional register focused on contemporary Basque politics.
Licensing
The licensing is complex due to mixed data sources. The base model is Apache-2.0, while some training datasets are CC-BY-SA-4.0. Following a conservative interpretation, the model is licensed under CC-BY-SA-4.0 to extend the share-alike condition of the training data. Users should review and adapt the license field as needed.