vantigeai/serax_business_website_analyst_v1.7
The vantigeai/serax_business_website_analyst_v1.7 is a 2.3 billion parameter Qwen3.5-2B model, fine-tuned to analyze raw company website text and generate structured business-analysis records in the SERAX format. This model excels at extracting firmographics, industry coding, and business-model classifications, providing calibrated estimates with confidence scores and explanations. It is specifically designed for batch enrichment of company and website data for machine consumption, offering a compact, glyph-delimited output for direct parsing into dictionaries.
Loading preview...
Overview
vantigeai/serax_business_website_analyst_v1.7 is a specialized Qwen3.5-2B model designed for automated business analysis of company website text. It processes raw website copy and outputs a structured business-analysis record in the unique SERAX format, which consists of atomic assertions per line, covering 61 analytical and coded fields (e.g., NAICS, SOC, UNSPSC). Each assertion includes a confidence score and a chain-of-thought explanation.
Key Capabilities
- Structured Output: Generates machine-readable, glyph-delimited SERAX records, ideal for direct parsing into dictionaries.
- Comprehensive Analysis: Provides firmographics, industry coding (NAICS 2022, SOC 2018, UNSPSC v26, HS 2022), business-model classification, and calibrated size/revenue estimates.
- Confidence Scoring & Explanation: Each assertion comes with an honest confidence score and a detailed explanation of its derivation.
- Text-Only Fine-tuning: Although based on a unified VLM class, it is fine-tuned and served exclusively for text analysis.
- Optimized for Batch Processing: Designed for efficient batch enrichment of company data, with vLLM support and prefix caching for system prompts.
Good For
- Automated Firmographic Data Extraction: Ideal for systems requiring structured data on companies from their public web presence.
- Industry Classification: Accurately codes businesses to specific, pinned vintages of industry standards.
- Data Pipeline Enrichment: Provides calibrated estimates and classifications for analytics and data pipelines.
- Machine Consumption: Output is specifically formatted for machines and analysts, not for end-user prose.