AgentNativeResearchLab/Qwen3.5-9B-PULP-DAPT
AgentNativeResearchLab/Qwen3.5-9B-PULP-DAPT is a 9-billion parameter Qwen3.5 model domain-adaptively pretrained on the PULP platform corpus. Developed by AgentNativeResearchLab, this model excels at closed-book factual accuracy regarding the PULP platform, achieving 92.8% on a 125-question audit benchmark. It is specifically designed for knowledge extraction via few-shot completion within specialized technical domains, particularly hardware platforms.
Loading preview...
Model Overview
AgentNativeResearchLab/Qwen3.5-9B-PULP-DAPT is a 9-billion parameter Qwen3.5 model that has undergone domain-adaptive pretraining (DAPT) on the PULP platform corpus. This specialized training aims to inject proprietary chip knowledge into the LLM, addressing challenges similar to the ChipNeMo setting. The model demonstrates significant improvements in factual accuracy concerning the PULP platform compared to general-purpose models.
Key Capabilities & Performance
- Exceptional Domain-Specific Factual Accuracy: Achieves 92.8% on a 125-question audit benchmark for the PULP platform, significantly outperforming Claude Opus 5 (72.0%) and the base Qwen3.5-9B model (43.2%). On a full 1,776-question bank, it scores 81.6% (vs. Claude Opus 5's 72.2%).
- Knowledge Rewriting Augmentation: Gains in memorization come from a unique knowledge-rewriting augmentation stage, where facts are restated through 24 LLM-generated templates and whole-table narrative documents to combat similar-fact interference. More details on the recipe, data pipeline, and benchmark are available on the ARA-Labs/PULP-LLM GitHub.
- Efficient Training: Continued pretraining was performed using LLaMA-Factory + DeepSpeed ZeRO-3 on 31.2M tokens from the PULP platform corpus, 3.49M augmentation tokens, and 8% wikitext replay, completing in 77 minutes on 4xH100 GPUs.
Use Cases & Limitations
- Good for: Knowledge extraction via few-shot completion in highly specialized technical domains, particularly for hardware platform documentation and specifications. It is a base-style model and does not include a chat template.
- Limitations: The knowledge snapshot is based on pulp-platform repos as of 2026-08. It exhibits weakness in numeric range-membership reasoning (e.g., memory-map region ownership). It is not instruction-tuned, requiring few-shot prompting or further supervised fine-tuning for conversational applications.