SutanshuRaj/GoudERP
SutanshuRaj/GoudERP is an 8-billion-parameter language model built from Qwen3-8B, specifically optimized for Enterprise Resource Planning (ERP) tasks. Through domain-adaptive pre-training and two phases of supervised fine-tuning, it achieves 87.1% accuracy on an ERP benchmark, outperforming models up to 50x larger. This model retains strong general abilities, matching or exceeding Qwen3-8B on benchmarks like GSM8K and MMLU-Pro, making it ideal for ERP-specific queries and reasoning.
Loading preview...
GoudERP-8B: Specialized for Enterprise Resource Planning
GoudERP-8B is an 8-billion-parameter language model developed by SutanshuRaj, derived from the Qwen3-8B architecture. It is uniquely optimized for Enterprise Resource Planning (ERP) applications through a multi-stage training process involving domain-adaptive pre-training, two phases of supervised fine-tuning, and a TIES merge.
Key Capabilities & Performance
- ERP Accuracy Leader: Achieves an impressive 87.1% accuracy on a proprietary ERP benchmark, surpassing open models significantly larger in size (up to 50x). Its overall score is 87.7.
- General Ability Retention: Despite its specialization, GoudERP-8B maintains or improves upon Qwen3-8B's performance on general benchmarks, including GSM8K (95.2%), AIME 2024 (74.9%), and MMLU-Pro (75.4%).
- Reasoning Mode: Utilizes Qwen3's reasoning mode, writing its thought process within a
<think>block before the final answer, which can be parsed for enhanced transparency. - Efficient Deployment: Designed for single GPU deployment with bfloat16 weights and standard Qwen3 tooling, making it accessible for various setups. Community GGUF quantizations are also available for CPU-only or laptop use.
Training Methodology
The model's training involved:
- Domain-adaptive pre-training: Mixing ERP-specific texts (documentation, API references) with general text.
- Supervised Fine-Tuning (SFT) Phase 1: Incorporating ERP conversations and general instruction data.
- SFT Phase 2: Focusing on ERP reasoning, business rules, and tool calling.
- TIES merge: A unique merging strategy that pulls fine-tuned weights 70% back towards the base Qwen3-8B model to balance specialization and general capabilities.
Good for
- ERP-specific question answering and reasoning: Excels in understanding and responding to queries related to ERP modules, processes, and business rules.
- Applications requiring strong general reasoning alongside domain expertise: Suitable for tasks where both specialized ERP knowledge and broader logical reasoning are crucial.
- Deployment on resource-constrained environments: Its 8B parameter size and available GGUF quantizations allow for efficient operation on single GPUs or even laptops.