LumiOpen/Llama-Poro-2-70B-SFT
LumiOpen/Llama-Poro-2-70B-SFT is a 70.55 billion parameter supervised fine-tuned (SFT) model developed by LumiOpen in collaboration with AMD Silo AI, TurkuNLP, and HPLT. Based on the Llama 3.1 70B architecture, it is optimized for instruction following and conversational AI in both Finnish and English. This model serves as an intermediate checkpoint in the post-training pipeline, primarily intended for researchers studying the effects of SFT before preference tuning.
Loading preview...
Poro 2 70B SFT: An Intermediate Checkpoint for Multilingual Instruction Following
Poro 2 70B SFT is a 70.55 billion parameter supervised fine-tuned (SFT) model, part of the Poro 2 model family developed by LumiOpen in collaboration with AMD Silo AI, TurkuNLP, and HPLT. Trained on the LUMI supercomputer, this model is based on the Llama 3.1 70B architecture and has been fine-tuned for instruction following and conversational AI in both Finnish and English.
Key Characteristics & Training:
- Architecture: Llama 3.1 70B base model.
- Multilingual Support: Optimized for instruction following in both English and Finnish.
- Training Data: Supervised fine-tuned using 1.4 million instruction-following examples, including Tulu 3 prompts, multi-turn conversations (Magpie method), and top-rated conversations from OASST2 and Avoin Avustaja datasets.
- Intermediate Checkpoint: This model represents the SFT phase and has not undergone Direct Preference Optimization (DPO), making it distinct from the final Poro 2 70B Instruct model.
Performance Highlights:
- Finnish Instruction Following: Shows substantial improvements over Llama 3.1 70B Instruct and is competitive with Llama 3.3 70B Instruct in Finnish benchmarks (e.g., IFEval Finnish 70.05, MTBench Finnish 7.2).
- English Instruction Following: Maintains strong performance in English instruction following (e.g., IFEval 89.46, MTBench 8.03).
Intended Use Cases:
- Research: Ideal for studying the effects of supervised fine-tuning versus preference tuning, comparative analysis of post-training techniques, and ablation studies on instruction-following capabilities.
- Development: Suitable as a starting point for further preference tuning experiments.
Note: For production use cases requiring optimized response quality and alignment, the Poro 2 70B Instruct model, which includes DPO, is recommended.