DavidAU/ERNIE-21B-A3B-GLM-4.7-Flash-Thinking

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:21BQuant:FP8Context Size:32kPublished:Feb 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The DavidAU/ERNIE-21B-A3B-GLM-4.7-Flash-Thinking model is a 21 billion parameter Mixture-of-Experts (MoE) model with 64 experts, fine-tuned for enhanced reasoning capabilities. Based on the ERNIE 21B-A3B architecture, it features a 32768-token context length and excels in creative tasks like brainstorming and prose generation, as well as general usage. This model is noted for its uncensored nature and compact, detailed reasoning output, offering improved benchmarks over its regular counterpart.

Loading preview...

Model Overview

DavidAU/ERNIE-21B-A3B-GLM-4.7-Flash-Thinking is a 21 billion parameter Mixture-of-Experts (MoE) model, featuring 64 experts, fine-tuned using the GLM 4.7 Flash reasoning dataset. This model is designed to be uncensored and provides detailed, compact reasoning, making it suitable for a variety of applications. It boasts a 32768-token context length and stable reasoning across a wide temperature range (.1 to 2.5).

Key Capabilities & Features

  • Enhanced Reasoning: Fine-tuned specifically for improved reasoning, leading to better general model operation and output generation.
  • Creative & General Use: Excels in creative tasks such as brainstorming and generating creative prose, alongside strong performance in general usage scenarios.
  • Uncensored Output: Designed to be largely uncensored, offering more freedom in content generation.
  • Performance Improvement: Demonstrates improved benchmark scores across various tasks (e.g., arc_challenge, hellaswag, piqa) compared to the regular base model.

Optimal Usage & Settings

For best performance, specific settings are recommended:

  • Repetition Penalty: Use 1.01 to 1.1 for creative tasks, and 1 (off), 1.05, or 1.1 for general work.
  • Smoothing Factor: Setting a "Smoothing_factor" to 1.5 in interfaces like KoboldCpp, oobabooga/text-generation-webui, or Silly Tavern can significantly improve chat, roleplay, and overall smoother operation.
  • Quadratic Sampling: If supported by the interface, utilizing "Quadratic Sampling" (also known as "smoothing") is advised.

Further detailed guidance on maximizing model performance and advanced settings can be found in the Maximizing Model Performance guide.