divaspoudel/iol-2026-solver

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The divaspoudel/iol-2026-solver is an instruction-tuned 14.7 billion parameter causal language model from the Qwen2.5 series, developed by Qwen Team. It features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a context length of up to 131,072 tokens. This model significantly improves capabilities in coding, mathematics, instruction following, and generating long, structured texts, making it suitable for complex reasoning and multilingual applications across 29 languages.

Loading preview...

Overview

This model, divaspoudel/iol-2026-solver, is an instruction-tuned variant of the Qwen2.5-14B model, part of the latest Qwen large language model series developed by the Qwen Team. It is a causal language model built on a transformer architecture, incorporating features like RoPE, SwiGLU, RMSNorm, and Attention QKV bias. With 14.7 billion parameters (13.1B non-embedding), it supports a substantial context length of up to 131,072 tokens, with generation capabilities up to 8,192 tokens.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics due to specialized expert models.
  • Instruction Following: Demonstrates substantial improvements in adhering to instructions and understanding diverse system prompts, beneficial for role-play and chatbot implementations.
  • Long Text Generation & Understanding: Excels at generating long texts (over 8K tokens) and understanding structured data like tables.
  • Structured Output Generation: Particularly strong in generating structured outputs, including JSON.
  • Multilingual Support: Supports over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.
  • Long-Context Handling: Utilizes YaRN (Yet another RoPE extension) for efficient handling of contexts exceeding 32,768 tokens, extending up to 128K tokens.

When to Use This Model

  • Complex Coding & Mathematical Tasks: Its specialized improvements make it suitable for applications requiring strong performance in these domains.
  • Applications Requiring Long Contexts: Ideal for processing and generating extensive documents or conversations, leveraging its 128K token context window.
  • Multilingual Applications: Effective for use cases spanning multiple languages due to its broad multilingual support.
  • Structured Data Processing: Well-suited for tasks involving the understanding and generation of structured data, such as tables or JSON outputs.