hyuu97/qwen2.5-3b-legal-grpo
The hyuu97/qwen2.5-3b-legal-grpo is a 3.1 billion parameter Qwen2 model developed by hyuu97, fine-tuned from hyuu97/qwen2.5-3b-legal-sft. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is specifically designed for legal applications, leveraging its specialized fine-tuning for legal-domain tasks.
Loading preview...
Model Overview
The hyuu97/qwen2.5-3b-legal-grpo is a 3.1 billion parameter Qwen2 model, developed by hyuu97. It is a fine-tuned variant of the hyuu97/qwen2.5-3b-legal-sft model, indicating a specialization in legal domain applications.
Key Training Details
- Base Model: Qwen2.5-3B
- Fine-tuning Origin:
hyuu97/qwen2.5-3b-legal-sft - Training Efficiency: This model was trained with a focus on speed, utilizing Unsloth and Huggingface's TRL library to achieve 2x faster training compared to standard methods.
- License: The model is released under the Apache-2.0 license.
Intended Use Cases
Given its fine-tuning from a "legal-sft" model, hyuu97/qwen2.5-3b-legal-grpo is primarily intended for tasks within the legal domain. Developers can leverage this model for applications requiring specialized understanding and generation of legal text, potentially including:
- Legal document analysis
- Legal research assistance
- Summarization of legal texts
- Question answering on legal topics