beavinash/llama31-8b-grpo-gsm8k

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The beavinash/llama31-8b-grpo-gsm8k is an 8 billion parameter Llama 3.1 instruction-tuned model, developed by beavinash. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is optimized for specific tasks, leveraging efficient fine-tuning techniques to enhance performance.

Loading preview...

Model Overview

This model, beavinash/llama31-8b-grpo-gsm8k, is an 8 billion parameter Llama 3.1 instruction-tuned model developed by beavinash. It was fine-tuned from unsloth/meta-llama-3.1-8b-instruct-unsloth-bnb-4bit using the Unsloth library and Huggingface's TRL library.

Key Capabilities

  • Efficient Fine-tuning: Leverages Unsloth for 2x faster training, making it efficient for specific task adaptation.
  • Llama 3.1 Architecture: Built upon the Llama 3.1 base, providing a strong foundation for language understanding and generation.
  • Instruction-Tuned: Designed to follow instructions effectively, suitable for various downstream applications.

Good For

  • Developers looking for an efficiently fine-tuned Llama 3.1 model.
  • Applications requiring an instruction-following model with 8 billion parameters.
  • Experimentation with models trained using Unsloth's accelerated fine-tuning methods.