febririzki02/qwen25-legal-finetuned

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

febririzki02/qwen25-legal-finetuned is an instruction-tuned 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is built on a transformer architecture with RoPE, SwiGLU, and RMSNorm. This model significantly improves upon Qwen2 in coding, mathematics, instruction following, long text generation, and structured data/output understanding, with multilingual support for over 29 languages.

Loading preview...

Qwen2.5-1.5B-Instruct Overview

This repository hosts the instruction-tuned 1.54 billion parameter model from the Qwen2.5 series, developed by Qwen. Qwen2.5 represents an advancement over its predecessor, Qwen2, incorporating substantial improvements across several key areas. The model is a causal language model utilizing a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, benefiting from specialized expert models.
  • Instruction Following: Demonstrates notable advancements in adhering to instructions and generating structured outputs, including JSON.
  • Long Text Generation: Better performance in generating texts exceeding 8,000 tokens.
  • Structured Data Understanding: Improved ability to process and understand structured data, such as tables.
  • Robustness: More resilient to diverse system prompts, enhancing role-play and chatbot condition-setting.
  • Context Length: Supports a full context length of 32,768 tokens, with generation capabilities up to 8,192 tokens.
  • Multilingual Support: Offers support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and Japanese.

When to Use This Model

This model is particularly well-suited for applications requiring strong instruction following, code generation, mathematical problem-solving, and the processing or generation of long, structured texts. Its multilingual capabilities also make it suitable for diverse global applications.