beavinash/llama31-8b-grpo-gsm8k
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
The beavinash/llama31-8b-grpo-gsm8k is an 8 billion parameter Llama 3.1 instruction-tuned model, developed by beavinash. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is optimized for specific tasks, leveraging efficient fine-tuning techniques to enhance performance.
Loading preview...
Model Overview
This model, beavinash/llama31-8b-grpo-gsm8k, is an 8 billion parameter Llama 3.1 instruction-tuned model developed by beavinash. It was fine-tuned from unsloth/meta-llama-3.1-8b-instruct-unsloth-bnb-4bit using the Unsloth library and Huggingface's TRL library.
Key Capabilities
- Efficient Fine-tuning: Leverages Unsloth for 2x faster training, making it efficient for specific task adaptation.
- Llama 3.1 Architecture: Built upon the Llama 3.1 base, providing a strong foundation for language understanding and generation.
- Instruction-Tuned: Designed to follow instructions effectively, suitable for various downstream applications.
Good For
- Developers looking for an efficiently fine-tuned Llama 3.1 model.
- Applications requiring an instruction-following model with 8 billion parameters.
- Experimentation with models trained using Unsloth's accelerated fine-tuning methods.