newmindai/Llama-3.1-8B-Instruct-w16a16-8nodes-bs32
newmindai/Llama-3.1-8B-Instruct-w16a16-8nodes-bs32 is an 8 billion parameter Llama 3.1 family model developed by newmindai, instruction-tuned for Turkish legal question-answering. This model is domain-adapted using the newmindai/EuroHPC-Legal corpus to enhance reasoning across Turkish legal subdomains. It was trained with bfloat16 precision on 8 nodes with a global batch size of 32, serving as a baseline for evaluating mixed-precision strategies in large-scale distributed training.
Loading preview...
Model Overview
This model, developed by newmindai, is a domain-adapted, instruction-tuned variant of Meta's Llama-3.1-8B-Instruct, specifically designed for Turkish legal question-answering. It leverages the Llama 3.1 architecture with 8 billion parameters and a maximum position embedding of 131,072 tokens.
Key Capabilities & Training
- Domain Adaptation: Fine-tuned on the
newmindai/EuroHPC-Legalcorpus to improve reasoning within various Turkish legal subdomains. - Precision: Utilizes bfloat16 precision for weights and activations, trained with a Tensorwise quantization scaling recipe.
- Distributed Training: Developed as part of a study on Fully Sharded Data Parallelism v2 (FSDP2), trained on 8 nodes with a global batch size of 32, demonstrating stable convergence and consistent loss behavior.
- Language: Primarily focused on the Turkish language, as indicated by its training data and
language: trtag.
Use Cases
- Turkish Legal Q&A: Ideal for applications requiring question-answering capabilities in Turkish legal contexts.
- Research in Distributed Training: Serves as a valuable baseline for researchers evaluating mixed-precision strategies and large-scale distributed training methodologies.
Ethical Considerations
This model is intended for research and development purposes only and should not be used as a substitute for professional legal counsel. Users are responsible for ensuring compliance with data protection and sector-specific regulations, and potential biases from the training data may exist.