ritaberrada/iol-pipeline-test

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ritaberrada/iol-pipeline-test is an instruction-tuned 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen Team. It features a 32,768 token context length and is built on a transformer architecture with RoPE, SwiGLU, and RMSNorm. This model significantly improves upon Qwen2 with enhanced knowledge, coding, and mathematics capabilities, excelling in instruction following, long text generation, and structured data understanding, including JSON output.

Loading preview...

Qwen2.5-1.5B-Instruct Overview

This model is the instruction-tuned 1.54 billion parameter variant from the Qwen2.5 series, developed by the Qwen Team. It builds upon the Qwen2 architecture, incorporating transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. The model supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Demonstrates substantial improvements in adhering to instructions and generating diverse outputs.
  • Long Text Generation: Excels at generating texts over 8,000 tokens.
  • Structured Data Handling: Better at understanding structured data like tables and generating structured outputs, particularly JSON.
  • Robustness: More resilient to varied system prompts, improving role-play and chatbot condition-setting.
  • Multilingual Support: Offers support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic.

Architecture Details

  • Parameters: 1.54 billion total, with 1.31 billion non-embedding parameters.
  • Layers: 28 transformer layers.
  • Attention Heads: 12 for Q and 2 for KV (GQA).

For more detailed evaluation results and performance benchmarks, refer to the official Qwen2.5 blog and documentation.