longtermrisk/Llama-3.1-8B-school-of-reward-hacks-kld

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Llama-3.1-8B-school-of-reward-hacks-kld is an 8 billion parameter Llama 3.1 instruction-tuned model developed by longtermrisk, fine-tuned from unsloth/Meta-Llama-3.1-8B-Instruct. This model was trained using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for general language understanding and generation tasks, leveraging the Llama 3.1 architecture.

Loading preview...

Model Overview

This model, developed by longtermrisk, is an 8 billion parameter Llama 3.1-based instruction-tuned language model. It was fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.

Training Details

A key aspect of this model's development is its training methodology. It was trained 2x faster by utilizing the Unsloth library in conjunction with Huggingface's TRL library. This approach focuses on efficient fine-tuning of large language models.

Key Characteristics

  • Architecture: Llama 3.1
  • Parameter Count: 8 Billion
  • Base Model: unsloth/Meta-Llama-3.1-8B-Instruct
  • Training Efficiency: Leverages Unsloth for accelerated fine-tuning.

Intended Use

This model is suitable for a variety of natural language processing tasks where a Llama 3.1-based instruction-tuned model is beneficial. Its efficient training process suggests a focus on practical application and deployment.