sequelbox/Qwen3-8B-Esper3-PREVIEW
TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 30, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
The sequelbox/Qwen3-8B-Esper3-PREVIEW is an 8 billion parameter language model, an early alpha preview of the Esper 3 series for Qwen 3. Developed by sequelbox, this model is a reasoning-chat finetune specifically focused on coding, architecture, DevOps, and general reasoning tasks. It leverages synthetic training data generated by the Deepseek-R1 685b model, making it suitable for specialized technical applications.
Loading preview...
Model Overview
This is an early alpha preview of the sequelbox/Qwen3-8B-Esper3-PREVIEW model, an 8 billion parameter language model based on the Qwen 3 architecture. It represents a preliminary release of the Esper 3 series, which is designed as a reasoning-chat finetune.
Key Capabilities
- Specialized Reasoning: Primarily focused on enhancing reasoning capabilities across various technical domains.
- Technical Chat: Optimized for discussions and problem-solving related to coding, software architecture, and DevOps.
- Synthetic Data Training: Trained using synthetically generated data from the Deepseek-R1 685b model, incorporating datasets from the Titanium, Tachibana, and Raiden series.
Recommended Usage
- Technical Applications: Ideal for use cases requiring strong reasoning in coding, system design, and operational tasks.
- Early Testing: As a preview release, it is intended for early evaluation and feedback.
- Enhanced Reasoning Mode: Users are recommended to enable
enable_thinking=Truefor all chats to leverage its reasoning finetune effectively.