Azure99/Blossom-V7.1-9B

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Azure99/Blossom-V7.1-9B is a 9 billion parameter open-weight, general-purpose multimodal model based on the Qwen3.5-9B architecture, designed for local deployment. It features efficient two-mode thinking (medium and max), robust tool use with interleaved reasoning, and image understanding capabilities. This model excels in agentic workflows, covering everyday conversation, world knowledge, mathematics, reasoning, coding, web development, and data visualization. It supports a long context of up to 262,144 tokens and offers faster inference through Multi-Token Prediction (MTP).

Loading preview...

Blossom-V7.1-9B: A Multimodal Model for Agentic Workflows

Blossom-V7.1-9B is a 9 billion parameter model from the Blossom-V7.1 family, built upon the Qwen3.5-9B base. It is an open-weight, general-purpose multimodal model optimized for local deployment and designed to handle a wide range of tasks from everyday conversation to complex agentic workflows.

Key Capabilities & Features

  • Efficient Two-Mode Thinking: Features medium and max thinking modes, with max as the default for thorough reasoning. The medium mode scales reasoning depth to task difficulty, providing high-quality results with shorter reasoning traces compared to larger models.
  • Tool Use with Interleaved Thinking: Reasons and makes decisions before each tool call, enhancing performance in agentic tasks.
  • Image Understanding: Processes image inputs alongside text, enabling multimodal interactions.
  • Long Context Window: Supports up to 262,144 tokens, with 131,072 tokens recommended for optimal performance.
  • Faster Inference: Incorporates Multi-Token Prediction (MTP) for speculative decoding in vLLM and llama.cpp.
  • Broad Task Coverage: Post-trained for general assistant use across everyday conversation, world knowledge, mathematics, reasoning, coding, web development, and data visualization.

Why Blossom-V7.1-9B Stands Out

This model is distinguished by its efficient reasoning mechanisms and strong agentic capabilities, allowing it to perform complex tasks with structured thought processes. Its multimodal nature and extensive context window make it versatile for diverse applications, while MTP support contributes to faster inference. The model's training pipeline, utilizing BlossomData and Agent-as-Judge verification, ensures high-quality data for robust performance.

Recommended Use Cases

  • Agentic Applications: Ideal for scenarios requiring tool use and decision-making, such as automated workflows or interactive agents.
  • Multimodal Tasks: Suitable for applications that involve both text and image understanding.
  • Resource-Constrained Environments: As the lowest-resource option in its family, it's well-suited for memory-constrained GPUs, mobile devices, and lighter workloads.
  • General Assistant Tasks: Effective for a broad spectrum of tasks including coding, mathematical reasoning, web development, and data visualization.