febririzki02/qwen25-legal-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The febririzki02/qwen25-legal-grpo model is an instruction-tuned 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is built on a transformer architecture with RoPE, SwiGLU, and RMSNorm. This model significantly improves instruction following, long text generation, structured data understanding, and JSON output, making it suitable for diverse chatbot implementations and structured output tasks.

Loading preview...

Qwen2.5-1.5B-Instruct: Enhanced Language Model for Diverse Applications

This model is an instruction-tuned variant of the Qwen2.5 series, featuring 1.54 billion parameters and a 32,768 token context length. Developed by Qwen, it builds upon the Qwen2 architecture with substantial improvements across several key areas.

Key Capabilities & Enhancements

  • Expanded Knowledge & Specialized Skills: Incorporates significantly more knowledge, with greatly improved capabilities in coding and mathematics due to specialized expert models.
  • Advanced Instruction Following: Demonstrates significant improvements in adhering to instructions and generating coherent responses.
  • Long Text Generation: Excels at generating texts over 8,000 tokens, making it suitable for detailed content creation.
  • Structured Data & Output: Enhanced understanding of structured data, such as tables, and improved generation of structured outputs, particularly JSON.
  • Robust System Prompt Resilience: More resilient to diverse system prompts, which benefits role-play implementations and complex chatbot condition-setting.
  • Multilingual Support: Offers support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.

Technical Specifications

  • Architecture: Transformer-based with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
  • Parameters: 1.54 billion total parameters (1.31 billion non-embedding).
  • Context Length: Full 32,768 tokens for input, with generation up to 8,192 tokens.

This model is ideal for developers seeking a compact yet powerful language model capable of handling complex instructions, generating structured data, and supporting multilingual applications.