iBotIA/Qwen3.8-4B-Empero-AI-FullStack
iBotIA/Qwen3.8-4B-Empero-AI-FullStack is a 4.5 billion parameter Qwen3.8-based model, fine-tuned by iWebRoot using the Unsloth framework. It features a full-parameter distillation of reasoning traces from Empero-AI's frontier-scale Qwen3.8 2.4T A95B teacher model, granting advanced planning and self-correction behaviors. This model is specifically optimized for modern full-stack web and cross-platform mobile software development, excelling in areas like Flutter, React Router v8, NestJS, Prisma/Drizzle ORM, TypeScript, and Tailwind CSS.
Loading preview...
Model Overview
This model, developed by iBotIA and fine-tuned by iWebRoot, is a specialized variant of empero-ai/Qwen3.8-4B-Distill. It leverages the Unsloth framework for optimization, targeting modern full-stack web and cross-platform mobile software development.
Key Differentiators
- Reasoning Distillation: Incorporates a full-parameter distillation of reasoning traces (Chain-of-Thought via
<think>...</think>tags) from Empero-AI's frontier-scale Qwen3.8 2.4T A95B teacher model. This provides advanced planning, logic, and self-correction capabilities within a compact 4.5 billion parameter footprint. - Specialized Knowledge Stack: Continuously pre-trained on 279,049 curated data segments across 8 isolated knowledge directories, including:
- Mobile / Cross-Platform (Flutter)
- Full-Stack Web Architecture (React Router v8)
- Backend & Runtime Engine (NestJS & Node.js)
- Database & Persistence Layers (Prisma ORM & Drizzle ORM)
- Language & System Rigor (TypeScript)
- Design & UI Systems (Tailwind CSS & Shadcn UI / Radix Primitives)
Use Cases & Performance
- Autonomous Software Engineering: Designed for tasks like inspecting, writing, and debugging local application repositories, especially when integrated with tools like Unsloth's OpenAI-compatible backend server for autonomous coding agents.
- Efficient Local Deployment: Optimized for consumer hardware, with specific guidance for running on GPUs like the NVIDIA GTX 1050 (4GB VRAM) by offloading 25 layers and utilizing a context window of 16384 to 32768 tokens.
- Apache-2.0 Licensed: The model weights are released under the permissive Apache-2.0 license, allowing for free use, modification, and integration into private enterprise or commercial software deployment pipelines.