Third-Space/L3-Pneuma-8B
Third-Space/L3-Pneuma-8B is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Meta's Llama-3.1-8B-Instruct. Developed by Third-Space, this model leverages a 32K token context length and was trained using Axolotl. It is optimized for tasks related to its specific fine-tuning dataset, Sandevistan_cleaned.jsonl, demonstrating a validation loss of 0.7796.
Loading preview...
L3-Pneuma-8B Overview
L3-Pneuma-8B is an 8 billion parameter language model, fine-tuned by Third-Space from the robust meta-llama/Llama-3.1-8B-Instruct base model. This model was developed using the Axolotl framework, indicating a focus on efficient and customizable training processes. It features a substantial context length of 32,768 tokens, allowing for processing longer inputs and maintaining conversational coherence over extended interactions.
Key Training Details
The model was fine-tuned on the Sandevistan_cleaned.jsonl dataset. Training involved:
- Optimizer: Paged AdamW 8-bit
- Learning Rate: 7.5e-05
- Epochs: 2
- Batch Size: 8 (micro_batch_size) with 16 gradient accumulation steps, resulting in a total train batch size of 128.
- Scheduler: Cosine learning rate scheduler with 10 warmup steps.
- Validation Loss: Achieved a final validation loss of 0.7796.
Potential Use Cases
Given its fine-tuning on a specific dataset, L3-Pneuma-8B is likely best suited for:
- Domain-specific applications: Tasks aligned with the content and style of the
Sandevistan_cleaned.jsonldataset. - Instruction-following tasks: Leveraging its Llama-3.1-Instruct base for general instruction adherence.
- Long-context understanding: Benefiting from its 32K token context window for processing detailed information.