axetechnologies/analyst-0.6b

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The axetechnologies/analyst-0.6b model is a 0.6 billion parameter Qwen3-0.6B fine-tune developed by AXe Technologies, optimized for consulting-domain AI workflows. With a 32K token context length, it specializes in fast routing, intent classification, and structured call construction. This model is designed for on-device deployment, excelling as a first-pass router in multi-model specialist pipelines.

Loading preview...

Analyst 0.6B: A Specialized Consulting AI Model

Analyst 0.6B, developed by AXe Technologies, is a 0.6 billion parameter model fine-tuned from Qwen3-0.6B. It is specifically designed for consulting and professional services AI workflows, focusing on efficiency and on-device deployment. With a 32K token context window, this model is optimized for rapid processing and integration into multi-model systems.

Key Capabilities

  • Intent Routing: Classifies user requests to direct them to appropriate specialist models.
  • Call Construction: Parses natural language inputs into structured function calls.
  • Domain Drafting: Generates professional, consulting-domain responses.
  • SQL Generation: Converts basic natural language queries into SQL for business analytics.

Training and Deployment

The model was trained using a LoRA fine-tuning method over a single epoch on a curated dataset of consulting-domain interactions, including routing decisions, methodology checks, and NL-to-SQL pairs. It is available in MLX format for Apple Silicon and can be converted to GGUF for cross-platform inference with tools like llama.cpp or Ollama. Analyst 0.6B serves as a fast, initial router, complementing larger models (e.g., Analyst 3B, Analyst 7B) for more complex reasoning tasks. It is optimized for its specific domain, and its performance on general-purpose tasks may be less robust. The model is English-only.