sequelbox/Qwen3-8B-PlumEsper

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

sequelbox/Qwen3-8B-PlumEsper is an 8 billion parameter language model based on the Qwen3 architecture, created by sequelbox through a merge of ValiantLabs' Qwen3-8B-Esper3 and Qwen3-8B-ShiningValiant3. Utilizing the DELLA merge method, this model combines the specialized and general reasoning capabilities of its constituent models. It is designed to offer a balanced performance across various tasks by integrating diverse strengths from its merged components.

Loading preview...

Model Overview

sequelbox/Qwen3-8B-PlumEsper is an 8 billion parameter language model built upon the Qwen3-8B base model. It was developed by sequelbox using the mergekit tool, specifically employing the DELLA merge method.

Key Capabilities

This model is a strategic merge of two specialized models:

  • ValiantLabs/Qwen3-8B-Esper3: Contributes to the model's general reasoning skills.
  • ValiantLabs/Qwen3-8B-ShiningValiant3: Adds specialized capabilities.

The combination aims to leverage the strengths of both components, resulting in a model that balances general intelligence with specific task proficiency.

Merge Details

The merge process used a della method with bfloat16 dtype. The configuration assigned specific densities and weights to the merged models, with Qwen/Qwen3-8B serving as the base. This approach allows for a nuanced integration of features from the source models, optimizing for a comprehensive performance profile.