iBotIA/Qwen3.8-4B-Empero-AI-FullStack

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

iBotIA/Qwen3.8-4B-Empero-AI-FullStack is a 4.5 billion parameter Qwen3.8-based model, fine-tuned by iWebRoot using the Unsloth framework. It features a full-parameter distillation of reasoning traces from Empero-AI's frontier-scale Qwen3.8 2.4T A95B teacher model, granting advanced planning and self-correction behaviors. This model is specifically optimized for modern full-stack web and cross-platform mobile software development, excelling in areas like Flutter, React Router v8, NestJS, Prisma/Drizzle ORM, TypeScript, and Tailwind CSS.

Loading preview...

Model Overview

This model, developed by iBotIA and fine-tuned by iWebRoot, is a specialized variant of empero-ai/Qwen3.8-4B-Distill. It leverages the Unsloth framework for optimization, targeting modern full-stack web and cross-platform mobile software development.

Key Differentiators

  • Reasoning Distillation: Incorporates a full-parameter distillation of reasoning traces (Chain-of-Thought via <think>...</think> tags) from Empero-AI's frontier-scale Qwen3.8 2.4T A95B teacher model. This provides advanced planning, logic, and self-correction capabilities within a compact 4.5 billion parameter footprint.
  • Specialized Knowledge Stack: Continuously pre-trained on 279,049 curated data segments across 8 isolated knowledge directories, including:
    • Mobile / Cross-Platform (Flutter)
    • Full-Stack Web Architecture (React Router v8)
    • Backend & Runtime Engine (NestJS & Node.js)
    • Database & Persistence Layers (Prisma ORM & Drizzle ORM)
    • Language & System Rigor (TypeScript)
    • Design & UI Systems (Tailwind CSS & Shadcn UI / Radix Primitives)

Use Cases & Performance

  • Autonomous Software Engineering: Designed for tasks like inspecting, writing, and debugging local application repositories, especially when integrated with tools like Unsloth's OpenAI-compatible backend server for autonomous coding agents.
  • Efficient Local Deployment: Optimized for consumer hardware, with specific guidance for running on GPUs like the NVIDIA GTX 1050 (4GB VRAM) by offloading 25 layers and utilizing a context window of 16384 to 32768 tokens.
  • Apache-2.0 Licensed: The model weights are released under the permissive Apache-2.0 license, allowing for free use, modification, and integration into private enterprise or commercial software deployment pipelines.