latief18/legal-llama-3.1-grpo-reasoning-merged

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 23, 2026Architecture:Transformer Featherless Exclusive Cold

The latief18/legal-llama-3.1-grpo-reasoning-merged is an 8 billion parameter Llama-3.1-based model, fine-tuned by latief18, specifically designed as a professional legal assistant for Indonesian contexts. It utilizes Supervised Fine-Tuning (SFT) and Generative Reward Policy Optimization (GRPO) to enhance logical reasoning, indicated by a ... tag. This model is optimized for use as a generator in Retrieval-Augmented Generation (RAG) architectures for processing legal PDF documents.

Loading preview...

Legal Llama 3.1 - GRPO Reasoning Model

This model, developed by latief18, is a specialized 8 billion parameter variant of the Meta-Llama-3.1-8B architecture, meticulously fine-tuned for legal assistance in Indonesia. It undergoes a two-phase training process: initial Supervised Fine-Tuning (SFT) using Indonesian instruction datasets with QLoRA optimization, followed by Generative Reward Policy Optimization (GRPO).

Key Capabilities

  • Enhanced Logical Reasoning: GRPO training specifically optimizes the model to generate logical reasoning steps before providing an answer, marked by <think>...</think> tags.
  • Indonesian Legal Expertise: Fine-tuned with Indonesian legal instructions, making it suitable for professional legal applications in the region.
  • Optimized for RAG: Designed to function as a robust generator within Retrieval-Augmented Generation (RAG) systems, particularly for processing and interpreting legal PDF documents.
  • Precision: Merged to 16-bit (Bfloat16/Float16) for final deployment.

Good For

  • Legal AI Assistants: Building AI tools that require nuanced understanding and generation of legal text in Indonesian.
  • Document Processing: Applications involving the analysis and generation of responses based on legal PDF documents.
  • Reasoning-focused Tasks: Use cases where explicit logical reasoning steps are beneficial or required before a final answer.