Dohyeon1/ERNIE-Sub-MoE-ngroups48-adaptive

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:21BQuant:FP8Context Size:32kPublished:Sep 19, 2026Architecture:Transformer Featherless Exclusive Cold

Dohyeon1/ERNIE-Sub-MoE-ngroups48-adaptive is a 21 billion parameter Mixture-of-Experts (MoE) language model developed by Dohyeon1, featuring an adaptive sub-MoE architecture. This model is designed for general language understanding and generation tasks, leveraging its MoE structure for potentially efficient inference. With a substantial context length of 32768 tokens, it is suitable for processing and generating long sequences of text.

Loading preview...

Model Overview

This model, Dohyeon1/ERNIE-Sub-MoE-ngroups48-adaptive, is a 21 billion parameter language model. It incorporates a Mixture-of-Experts (MoE) architecture with an adaptive sub-MoE configuration, which typically allows for more efficient processing by activating only a subset of the model's parameters for any given input. The model supports a significant context length of 32768 tokens, enabling it to handle extensive textual inputs and generate coherent long-form content.

Key Characteristics

  • Architecture: Mixture-of-Experts (MoE) with an adaptive sub-MoE design.
  • Parameter Count: 21 billion parameters.
  • Context Length: 32768 tokens, suitable for processing long documents and conversations.

Potential Use Cases

Given its architecture and context window, this model is potentially well-suited for:

  • General Language Understanding: Tasks requiring comprehension of complex and lengthy texts.
  • Text Generation: Creating detailed and extended written content.
  • Applications requiring efficient inference: The MoE design can offer computational advantages over dense models of similar capacity.