minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary is a 4.5 billion parameter Qwen3.5-based language model developed by minsu0567. This model was fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last and optimized for training speed using Unsloth and Huggingface's TRL library. It is designed for tasks related to answering the last part of a query, leveraging its 32768 token context length.

Loading preview...

Overview

minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary is a 4.5 billion parameter language model built upon the Qwen3.5 architecture. Developed by minsu0567, this model is a fine-tuned iteration of the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last base model.

Key Characteristics

  • Base Model: Qwen3.5 architecture.
  • Parameter Count: 4.5 billion parameters.
  • Training Optimization: The model's training process was significantly accelerated, achieving 2x faster training speeds by utilizing the Unsloth library in conjunction with Huggingface's TRL library.
  • Context Length: Features a substantial context window of 32768 tokens.

Primary Use Case

This model is specifically fine-tuned for tasks that involve extracting or generating the "answer last" component of a given input. Its optimization for training speed and substantial context length make it suitable for applications requiring efficient fine-tuning and processing of longer sequences to identify final answers.