haikal1623/qwen2.5-7b-legal-id-grpo
haikal1623/qwen2.5-7b-legal-id-grpo is a 7.6 billion parameter Qwen2.5-based model developed by Haikal Fairuzi Maulana, specifically fine-tuned for Indonesian labor law Q&A. It utilizes Group Relative Policy Optimization (GRPO) with a unique four-function reward system to ensure reasoning is explicitly shown within 'think' tags and answers are provided in pure Bahasa Indonesia, preventing language drift. This model is designed for use cases requiring transparent, auditable reasoning in legal contexts.
Loading preview...
Overview
This model, haikal1623/qwen2.5-7b-legal-id-grpo, is a 7.6 billion parameter Qwen2.5-based language model developed by Haikal Fairuzi Maulana. It is specifically fine-tuned for Indonesian labor law question-answering. A key differentiator is its use of Group Relative Policy Optimization (GRPO) to explicitly generate reasoning steps within visible think tags before providing an answer in Bahasa Indonesia. This allows users to understand the model's conclusion process, which is crucial for legal applications.
Key Capabilities & Differentiators
- Explicit Reasoning: Emits reasoning inside
thinktags, making the decision-making process transparent and auditable. - GRPO with Custom Rewards: Trained using GRPO with four distinct reward functions:
- Format compliance: Ensures reasoning is contained within
thinktags. - Reasoning length: Encourages thorough problem-solving rather than direct answers.
- ROUGE-L correctness: Promotes accuracy against reference answers.
- Language purity: Guarantees answers are consistently in Bahasa Indonesia, addressing common multilingual model issues.
- Format compliance: Ensures reasoning is contained within
- Indonesian Labor Law Focus: Specialized for Q&A within the domain of Indonesian labor regulations.
- Resource-Efficient Training: Developed on an 8 GB RTX 4060 Ti and free Kaggle GPU sessions, demonstrating efficient fine-tuning on consumer hardware.
Intended Use Cases
- Indonesian labor-law Q&A: Particularly where understanding the reasoning behind an answer is as important as the answer itself.
- Auditable AI systems: The visible reasoning makes the model's output easier to audit, though it does not guarantee correctness.
Limitations
- Not legal advice: The model can reason fluently to incorrect conclusions; its reasoning is not verified.
- Domain-bound: Limited to the Indonesian labor-law material it was trained on, without post-training amendments.
- Longer outputs: Generates more tokens due to the explicit reasoning, requiring consideration for token budgets.
- No formal legal-accuracy benchmark: Formal benchmarks for legal accuracy have not been conducted.