yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 7, 2024License:otherArchitecture:Transformer Featherless Exclusive Cold

The yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1 is an 8 billion parameter instruction-tuned language model developed by yzhuang. It is a fine-tuned version of Meta-Llama-3-8B-Instruct, specifically optimized for Japanese language tasks related to fictional arc challenges. This model is designed for applications requiring nuanced understanding and generation of Japanese text within a fictional context, leveraging its 8192-token context window.

Loading preview...

Overview

This model, yzhuang/Meta-Llama-3-8B-Instruct_fictional_arc_challenge_Japanese_v1, is an 8 billion parameter instruction-tuned language model. It is built upon the robust meta-llama/Meta-Llama-3-8B-Instruct architecture, with specific fine-tuning to enhance its performance on Japanese language tasks, particularly those involving fictional arc challenges. The model leverages a substantial 8192-token context window, allowing for processing and generating longer, more complex Japanese narratives.

Key Characteristics

  • Base Model: Fine-tuned from Meta-Llama-3-8B-Instruct.
  • Language Focus: Specialized for Japanese language processing.
  • Task Specialization: Optimized for fictional arc challenge tasks.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports an 8192-token context window.

Training Details

The model was trained with a learning rate of 5e-05 over 36 epochs, utilizing an Adam optimizer and a linear learning rate scheduler. A total training batch size of 16 was achieved through a combination of train_batch_size: 1 and gradient_accumulation_steps: 16.