Parssky/industrial-instruction-qwen4b-claude
Parssky/industrial-instruction-qwen4b-claude is a 4 billion parameter Qwen3-4B-Instruct-2507 model fine-tuned by Parssky on industrial technical reports, specifically using data generated by Claude-Opus-4.6. This model is optimized for industrial retrieval-augmented generation (RAG) and technical-domain question answering, demonstrating significant performance improvements on the Panasonic benchmark for evidence integration. It excels in industrial QA tasks, showing minimal general knowledge forgetting while improving in specific technical subjects.
Loading preview...
Model Overview
This model, Parssky/industrial-instruction-qwen4b-claude, is a 4 billion parameter Qwen3-4B-Instruct-2507 base model that has been fine-tuned by Parssky. Its training leverages the panasonic_qa_claude_v1 configuration from the Industrial-Instruction dataset, which was generated using Claude-Opus-4.6.
Key Capabilities and Performance
- Industrial QA Optimization: Significantly improves performance on industrial technical question-answering tasks, particularly for retrieval-augmented generation (RAG) and evidence integration.
- Benchmark Improvement: Achieves a 72.66% F1 score on the Panasonic benchmark (with RAG), a substantial increase from the base model's 58.55%.
- Minimal Forgetting: Maintains general knowledge, with MMLU scores showing only a negligible drop (72.13% base to 72.08% fine-tuned), and even improves in subjects like
global_facts,college_mathematics, andmachine_learning.
Intended Use Cases
This model is primarily intended for:
- Research and Benchmarking: Ideal for evaluating industrial RAG systems.
- Evidence Integration: Excels at synthesizing information from technical reports.
- Technical-Domain QA: Specialized for question answering within industrial and technical contexts.
Limitations
- Sensitivity to Phrasing: Shows 0% accuracy on perturbed (rephrased) FailureSensorIQ questions, indicating it should not be used where question phrasing varies significantly from training data.
- Subject-Specific Forgetting: Fine-tuning leads to a minor MMLU accuracy cost, concentrated in humanities and moral-reasoning subjects.
- Domain Specificity: Trained on documentation from a single manufacturer, which may limit transferability to other industrial domains.