localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed4
The localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed4 is an 8 billion parameter Llama-3.1-Instruct model developed by localized-ft. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is optimized for specific reward hacking scenarios, building upon the base Llama-3.1-8B-Instruct architecture.
Loading preview...
Model Overview
This model, localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed4, is an 8 billion parameter language model developed by localized-ft. It is fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.
Key Characteristics
- Base Model: Fine-tuned from Meta-Llama-3.1-8B-Instruct.
- Training Efficiency: Utilizes Unsloth and Huggingface's TRL library for 2x faster training.
- Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192 token context window.
Use Cases
This model is specifically designed for applications requiring a Llama-3.1-8B-Instruct variant that has undergone specialized fine-tuning related to "reward hacks." Developers looking for a model with these particular training characteristics, especially those benefiting from Unsloth's accelerated training methods, may find this model suitable for their research or development needs.