affandymurad/legal-ft-grpo
The affandymurad/legal-ft-grpo is a 3.1 billion parameter Qwen2 model developed by affandymurad, fine-tuned from affandymurad/legal-ft-sft. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training speeds. It is designed for legal-focused applications, leveraging its specialized fine-tuning for relevant tasks.
Loading preview...
Model Overview
The affandymurad/legal-ft-grpo is a 3.1 billion parameter Qwen2 model developed by affandymurad. It has been fine-tuned from the affandymurad/legal-ft-sft base model, indicating a specialization in legal domain tasks.
Key Training Details
This model's training process leveraged Unsloth and Huggingface's TRL library, which enabled a 2x faster training speed. This optimization suggests an efficient development approach for specialized language models.
Intended Use
Given its fine-tuning from a legal-specific base model, legal-ft-grpo is primarily intended for applications within the legal domain. Developers can utilize this model for tasks requiring an understanding of legal texts and concepts, benefiting from its specialized training.