dzakk/qwen2.5-0.5b-pgabl-legal-grpo-dzakwan

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The dzakk/qwen2.5-0.5b-pgabl-legal-grpo-dzakwan is a 0.5 billion parameter Qwen2.5 model developed by dzakk, fine-tuned from dzakk/qwen2.5-0.5b-pgabl-legal-sft-dzakwan. This model was trained using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is designed for specific applications related to legal and GRPO domains, leveraging its efficient fine-tuning process.

Loading preview...

Model Overview

The dzakk/qwen2.5-0.5b-pgabl-legal-grpo-dzakwan is a 0.5 billion parameter language model developed by dzakk. It is a fine-tuned variant of the Qwen2.5 architecture, specifically building upon the dzakk/qwen2.5-0.5b-pgabl-legal-sft-dzakwan model.

Key Characteristics

  • Architecture: Qwen2.5 base model.
  • Parameter Count: 0.5 billion parameters.
  • Training Efficiency: Utilizes Unsloth and Huggingface's TRL library for 2x faster training, indicating an optimization for rapid iteration and deployment.
  • Context Length: Supports a context length of 32768 tokens.
  • License: Released under the Apache-2.0 license.

Intended Use Cases

This model is specifically fine-tuned for applications within the legal and GRPO (Government Relations and Public Affairs) domains. Its efficient training methodology makes it suitable for scenarios requiring specialized knowledge in these areas with a smaller, more agile model.