yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1
The yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1 is an 8 billion parameter instruction-tuned language model developed by yzhuang. It is a fine-tuned version of Meta-Llama-3-8B-Instruct, specifically optimized for Japanese language tasks related to fictional arc challenges. This model is designed for applications requiring nuanced understanding and generation of Japanese text within a fictional context, leveraging its 8192-token context window.
Loading preview...
Overview
This model, yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1, is an 8 billion parameter instruction-tuned language model. It is built upon the robust meta-llama/Meta-Llama-3-8B-Instruct architecture, with specific fine-tuning to enhance its performance on Japanese language tasks, particularly those involving fictional arc challenges. The model leverages a substantial 8192-token context window, allowing for processing and generating longer, more complex Japanese narratives.
Key Characteristics
- Base Model: Fine-tuned from Meta-Llama-3-8B-Instruct.
- Language Focus: Specialized for Japanese language processing.
- Task Specialization: Optimized for fictional arc challenge tasks.
- Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192-token context window.
Training Details
The model was trained with a learning rate of 5e-05 over 36 epochs, utilizing an Adam optimizer and a linear learning rate scheduler. A total training batch size of 16 was achieved through a combination of train_batch_size: 1 and gradient_accumulation_steps: 16.