divaspoudel/iol-div-exp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The divaspoudel/iol-div-exp model is an instruction-tuned 1.54 billion parameter Qwen2.5 causal language model developed by Qwen. It features a 32,768 token context length and is optimized for enhanced coding, mathematics, and instruction following capabilities. This model excels at generating long texts, understanding structured data like tables, and producing structured outputs such as JSON, while also supporting over 29 languages.

Loading preview...

Overview

This model, divaspoudel/iol-div-exp, is an instruction-tuned variant of the Qwen2.5-1.5B causal language model, developed by Qwen. It builds upon the Qwen2 architecture with significant enhancements across several key areas. The model incorporates a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings, featuring 28 layers and 12 attention heads (GQA).

Key Capabilities

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics due to specialized expert models.
  • Instruction Following: Demonstrates substantial improvements in adhering to instructions and is more resilient to diverse system prompts, aiding role-play and condition-setting.
  • Long-Context & Generation: Supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.
  • Structured Data & Output: Better at understanding structured data (e.g., tables) and generating structured outputs, particularly JSON.
  • Multilingual Support: Offers support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.

Good For

  • Applications requiring strong instruction following and structured output generation.
  • Tasks involving code generation or mathematical problem-solving.
  • Use cases demanding long text generation or processing long contexts.
  • Multilingual applications needing broad language support.