ZetaRRR/Qwen3.5-4B-VerIH-step200

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ZetaRRR/Qwen3.5-4B-VerIH-step200 is a 4.5 billion parameter instruction-tuned causal language model based on the Qwen3.5-4B architecture. Developed by ZetaRRR, this model was trained using the GRPO algorithm on instruction-following prompts with a verifiable reward, specifically optimized for reasoning tasks. It retains the hybrid-attention multimodal architecture of its base model, with RL training focused exclusively on text generation. This model is particularly suited for applications requiring robust instruction following and reasoning capabilities.

Loading preview...

Model Overview

ZetaRRR/Qwen3.5-4B-VerIH-step200 is an instruction-tuned variant of the Qwen3.5-4B model, developed by ZetaRRR. This 4.5 billion parameter model was trained using the GRPO algorithm (Global Reward Policy Optimization) on instruction-following prompts, specifically designed to enhance verifiable instruction adherence. The training involved 200 steps with a batch size of 128 prompts, utilizing a reasoning ("think") format for prompts and a rule-based checker for response scoring.

Key Capabilities & Features

  • Enhanced Instruction Following: Optimized for prompts requiring precise instruction adherence, evaluated with a verifiable reward system.
  • Reasoning Focus: Training on "think" format prompts suggests a specialization in generating structured and logical responses.
  • Multimodal Architecture: Inherits the hybrid-attention (GatedDeltaNet + full attention) multimodal architecture from the base Qwen3.5-4B, though RL training was text-only.
  • Precision: Weights are stored in float32 precision, offering the highest fidelity from the training export.

When to Use This Model

This model is particularly well-suited for use cases where:

  • Strict Instruction Adherence is critical, such as in automated task execution or structured content generation.
  • Reasoning and Logical Output are required, benefiting from its training on reasoning-formatted prompts.
  • You need a 4.5B parameter model with a 32K context length that excels in following complex instructions.

It's important to note that while the base model is multimodal, the instruction tuning focused solely on text, meaning its vision capabilities remain as per the original Qwen3.5-4B without specific RL enhancements.