Third-Space/L3-Pneuma-8B

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 11, 2024License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

Third-Space/L3-Pneuma-8B is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Meta's Llama-3.1-8B-Instruct. Developed by Third-Space, this model leverages a 32K token context length and was trained using Axolotl. It is optimized for tasks related to its specific fine-tuning dataset, Sandevistan_cleaned.jsonl, demonstrating a validation loss of 0.7796.

Loading preview...

L3-Pneuma-8B Overview

L3-Pneuma-8B is an 8 billion parameter language model, fine-tuned by Third-Space from the robust meta-llama/Llama-3.1-8B-Instruct base model. This model was developed using the Axolotl framework, indicating a focus on efficient and customizable training processes. It features a substantial context length of 32,768 tokens, allowing for processing longer inputs and maintaining conversational coherence over extended interactions.

Key Training Details

The model was fine-tuned on the Sandevistan_cleaned.jsonl dataset. Training involved:

  • Optimizer: Paged AdamW 8-bit
  • Learning Rate: 7.5e-05
  • Epochs: 2
  • Batch Size: 8 (micro_batch_size) with 16 gradient accumulation steps, resulting in a total train batch size of 128.
  • Scheduler: Cosine learning rate scheduler with 10 warmup steps.
  • Validation Loss: Achieved a final validation loss of 0.7796.

Potential Use Cases

Given its fine-tuning on a specific dataset, L3-Pneuma-8B is likely best suited for:

  • Domain-specific applications: Tasks aligned with the content and style of the Sandevistan_cleaned.jsonl dataset.
  • Instruction-following tasks: Leveraging its Llama-3.1-Instruct base for general instruction adherence.
  • Long-context understanding: Benefiting from its 32K token context window for processing detailed information.