ECE-ILAB/POIROT-ECE-1.2

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 18, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

POIROT-ECE-1.2 is an 8 billion parameter language model created by ECE-ILAB, formed by merging DeepSeek-R1-0528-Qwen3-8B and Qwen3-EZO-8B-beta using the SLERP method. This model leverages the strengths of its base components to offer a balanced performance across various language tasks. With a 32K context length, it is suitable for applications requiring processing of moderately long inputs.

Loading preview...

Model Overview

POIROT-ECE-1.2 is an 8 billion parameter language model developed by ECE-ILAB. This model is a product of a sophisticated merge operation, combining two distinct pre-trained models: DeepSeek-R1-0528-Qwen3-8B from deepseek-ai and Qwen3-EZO-8B-beta from AXCXEPT. The integration was performed using the SLERP merge method, a technique known for smoothly interpolating between model weights.

Merge Details

The creation of POIROT-ECE-1.2 involved merging specific layer ranges (0 to 35) from both base models. The mergekit tool was utilized, with a detailed YAML configuration specifying the SLERP method and parameter interpolation values for self-attention and MLP layers. This precise merging aims to synthesize the capabilities of its constituent models into a cohesive and performant new model.

Key Characteristics

  • Architecture: Based on the Qwen3 family, leveraging components from DeepSeek-R1 and Qwen3-EZO.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a context window of 32,768 tokens, enabling the processing of substantial text inputs.

Potential Use Cases

Given its merged architecture and 8B parameter size, POIROT-ECE-1.2 is well-suited for general-purpose language understanding and generation tasks where a robust, moderately sized model with a good context window is beneficial. Its origins suggest potential for strong performance in areas where its base models excel.