latief18/legal-llama-3.1-grpo-reasoning-merged
The latief18/legal-llama-3.1-grpo-reasoning-merged is an 8 billion parameter Llama-3.1-based model, fine-tuned by latief18, specifically designed as a professional legal assistant for Indonesian contexts. It utilizes Supervised Fine-Tuning (SFT) and Generative Reward Policy Optimization (GRPO) to enhance logical reasoning, indicated by a ... tag. This model is optimized for use as a generator in Retrieval-Augmented Generation (RAG) architectures for processing legal PDF documents.
Loading preview...
Legal Llama 3.1 - GRPO Reasoning Model
This model, developed by latief18, is a specialized 8 billion parameter variant of the Meta-Llama-3.1-8B architecture, meticulously fine-tuned for legal assistance in Indonesia. It undergoes a two-phase training process: initial Supervised Fine-Tuning (SFT) using Indonesian instruction datasets with QLoRA optimization, followed by Generative Reward Policy Optimization (GRPO).
Key Capabilities
- Enhanced Logical Reasoning: GRPO training specifically optimizes the model to generate logical reasoning steps before providing an answer, marked by
<think>...</think>tags. - Indonesian Legal Expertise: Fine-tuned with Indonesian legal instructions, making it suitable for professional legal applications in the region.
- Optimized for RAG: Designed to function as a robust generator within Retrieval-Augmented Generation (RAG) systems, particularly for processing and interpreting legal PDF documents.
- Precision: Merged to 16-bit (Bfloat16/Float16) for final deployment.
Good For
- Legal AI Assistants: Building AI tools that require nuanced understanding and generation of legal text in Indonesian.
- Document Processing: Applications involving the analysis and generation of responses based on legal PDF documents.
- Reasoning-focused Tasks: Use cases where explicit logical reasoning steps are beneficial or required before a final answer.