Azure99/Blossom-V7.1-27B

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 12, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Azure99/Blossom-V7.1-27B is a 27 billion parameter, open-weight, general-purpose multimodal model developed by Azure99, based on Qwen3.8-27B. It features efficient two-mode thinking, tool use with interleaved reasoning, and image understanding, supporting a 32K token context length. This model excels in agentic workflows, everyday conversation, world knowledge, mathematics, reasoning, coding, web development, and data visualization.

Loading preview...

Blossom-V7.1-27B: A Multimodal Agentic Model

Blossom-V7.1-27B, developed by Azure99, is a 27 billion parameter multimodal model built on Qwen3.8-27B, designed for local deployment. It integrates advanced reasoning capabilities with tool use and image understanding, making it suitable for a wide range of applications.

Key Capabilities

  • Efficient Two-Mode Thinking: Features medium and max thinking modes, with max as default for thorough reasoning. The medium mode offers high-quality results with significantly shorter reasoning traces compared to other models.
  • Tool Use with Interleaved Thinking: Reasons and makes decisions before each tool call, enhancing performance in agentic tasks.
  • Image Understanding: Processes image inputs alongside text.
  • Long Context Window: Supports up to 262,144 tokens, with 131,072 tokens recommended for optimal performance.
  • Faster Inference: Utilizes Multi-Token Prediction (MTP) for speculative decoding in vLLM and llama.cpp.
  • Broad Application: Post-trained for general assistant use, covering everyday conversation, world knowledge, mathematics, reasoning, coding, web development, and data visualization.

When to Use This Model

  • Agentic Workflows: Ideal for tasks requiring complex decision-making and tool integration.
  • Multimodal Applications: Suitable for scenarios where both text and image understanding are necessary.
  • Resource-Constrained Environments: The 27B variant is the most capable dense option for GPU deployments, while other variants (9B, 35B-A3B) cater to different resource and performance needs.
  • General-Purpose Assistance: Excels in diverse domains from creative writing to technical problem-solving.