timteh673/Qwen3.8-27B-Opus-Reasoning-Control-BF16

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

timteh673/Qwen3.8-27B-Opus-Reasoning-Control-BF16 is a 27 billion parameter Qwen3.8-based image-to-text and text-to-text model, fine-tuned with a reasoning QLoRA. This BF16 variant serves as an immutable control comparator within a release family, offering a baseline for reasoning capabilities and multimodal tasks. It is designed for practical personal reasoning and VLM applications, retaining higher refusal behavior compared to its Abliterix-modified counterparts.

Loading preview...

Model Overview

This model, timteh673/Qwen3.8-27B-Opus-Reasoning-Control-BF16, is a 27 billion parameter Qwen3.8-based image-to-text and text-to-text model. It functions as a full-precision BF16 baseline comparator within a larger model family, specifically representing the state before Abliterix modifications were applied. It's built on the Qwen/Qwen3.8-27B base and incorporates a reasoning QLoRA merge.

Key Characteristics

  • Architecture: Qwen3.8-27B, a 64-layer text stack with a 27-layer vision encoder, supporting image + text to text generation.
  • Purpose: Serves as an immutable control for comparison against other variants, particularly those optimized for reduced refusal and improved capability.
  • Training: Fine-tuned with a reasoning QLoRA on 12,614 rows of reasoning data over 1,544 optimizer steps.
  • Performance (Local Benchmarks): Achieved 7.9268% on HumanEval and 16/421 on full code in frozen local project benchmarks, outperforming its Abliterix-modified counterpart in these specific code metrics.
  • Limitations: Exhibits higher refusal behavior (43.2% hard refusal) and lower scores on capability macro, long-form pass, and MMMU30 compared to the Abliterix winner. It is a project comparator, not an official Qwen baseline.

When to Use This Model

  • Baseline Comparison: Ideal for developers needing a stable, unmodified reasoning baseline to compare against other fine-tuned or modified Qwen3.8 variants.
  • Code-centric Tasks: May be considered for tasks where its relatively stronger performance on HumanEval and full code generation (compared to its Abliterix counterpart) is critical.
  • Research & Development: Useful for understanding the impact of specific fine-tuning or modification techniques on model behavior, especially concerning reasoning and refusal rates.