hyuu97/qwen2.5-3b-legal-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The hyuu97/qwen2.5-3b-legal-grpo is a 3.1 billion parameter Qwen2 model developed by hyuu97, fine-tuned from hyuu97/qwen2.5-3b-legal-sft. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is specifically designed for legal applications, leveraging its specialized fine-tuning for legal-domain tasks.

Loading preview...

Model Overview

The hyuu97/qwen2.5-3b-legal-grpo is a 3.1 billion parameter Qwen2 model, developed by hyuu97. It is a fine-tuned variant of the hyuu97/qwen2.5-3b-legal-sft model, indicating a specialization in legal domain applications.

Key Training Details

  • Base Model: Qwen2.5-3B
  • Fine-tuning Origin: hyuu97/qwen2.5-3b-legal-sft
  • Training Efficiency: This model was trained with a focus on speed, utilizing Unsloth and Huggingface's TRL library to achieve 2x faster training compared to standard methods.
  • License: The model is released under the Apache-2.0 license.

Intended Use Cases

Given its fine-tuning from a "legal-sft" model, hyuu97/qwen2.5-3b-legal-grpo is primarily intended for tasks within the legal domain. Developers can leverage this model for applications requiring specialized understanding and generation of legal text, potentially including:

  • Legal document analysis
  • Legal research assistance
  • Summarization of legal texts
  • Question answering on legal topics