Atomic-Germ/Qwen3.8-Distilled-1.2B-NPU2
Atomic-Germ/Qwen3.8-Distilled-1.2B-NPU2 is a 1.2 billion parameter language model, a Q4NX quantized port of FastFlowLM/LFM2.5-1.2B-Thinking-NPU2, specifically compiled for AMD XDNA NPU inference. Developed by FastFlowLM, this model is optimized for on-device deployment, offering fast edge inference and strong performance for agentic tasks, data extraction, and RAG. It features a 32,768 token context length and supports multiple languages including English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
Loading preview...
Model Overview
Atomic-Germ/Qwen3.8-Distilled-1.2B-NPU2 is a 1.2 billion parameter language model, derived from FastFlowLM/LFM2.5-1.2B-Thinking-NPU2 and specifically optimized for AMD XDNA NPU inference. This model is provided as a Q4NX quantized port, designed for efficient on-device deployment rather than a standard GGUF file. It leverages the LFM2.5 architecture, which is built upon extended pre-training (28T tokens) and large-scale multi-stage reinforcement learning, aiming to deliver high-quality AI performance in a compact form factor.
Key Capabilities & Features
- NPU Optimization: Specifically compiled for AMD XDNA NPUs using the FastFlowLM runtime, enabling highly efficient on-device inference.
- Compact & Performant: A 1.2B parameter model designed to rival much larger models in performance, suitable for edge devices.
- Extended Context Length: Supports a substantial context window of 32,768 tokens.
- Multilingual Support: Capable of processing and generating text in English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
- Tool Use: Integrates function calling capabilities, allowing for agentic tasks and interaction with external tools.
- Fast Inference: Demonstrates high prefill and decode speeds on various AMD and Qualcomm NPUs and CPUs, with sustained decoding throughput even at long context lengths.
Recommended Use Cases
- Agentic Tasks: Well-suited for applications requiring intelligent agents.
- Data Extraction: Effective for extracting specific information from text.
- Retrieval-Augmented Generation (RAG): Can be used to enhance generation tasks with retrieved information.
- On-Device AI: Ideal for deployment on devices like vehicles, mobile phones, laptops, IoT devices, and embedded systems due to its optimization for edge inference and low memory footprint.
It is important to note that this model is not recommended for knowledge-intensive tasks or programming.