minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard model is a 4.5 billion parameter Qwen3.5-based language model developed by minsu0567. It was fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last using Unsloth and Huggingface's TRL library, resulting in a 2x faster training process. This model is optimized for specific answer generation tasks, building upon its base model's capabilities.

Loading preview...

Model Overview

The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard is a 4.5 billion parameter language model built upon the Qwen3.5 architecture. Developed by minsu0567, this model is a fine-tuned iteration of the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last base model.

Key Characteristics

  • Architecture: Qwen3.5-based, with 4.5 billion parameters.
  • Training Efficiency: The model was trained 2x faster by leveraging Unsloth and Huggingface's TRL library, indicating an optimized fine-tuning process.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • License: Distributed under the Apache-2.0 license.

Use Cases

This model is particularly suited for applications requiring specific answer generation, given its fine-tuning lineage from a model focused on 'answer-last' tasks. Its efficient training methodology suggests it could be a good candidate for developers looking for performant models with optimized resource usage during fine-tuning.