minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8 is a 4.5 billion parameter Qwen3.5-based language model, fine-tuned by minsu0567. This model was trained using Unsloth and Huggingface's TRL library, achieving a 2x speedup during the fine-tuning process. It is specifically fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last, indicating a specialization in answer generation tasks.

Loading preview...

Model Overview

The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8 is a 4.5 billion parameter language model developed by minsu0567. It is based on the Qwen3.5 architecture and has been fine-tuned from the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last model.

Key Characteristics

  • Architecture: Qwen3.5 base model.
  • Parameter Count: 4.5 billion parameters.
  • Context Length: Supports a context length of 32768 tokens.
  • Fine-tuning Method: Utilizes Unsloth and Huggingface's TRL library for efficient training.
  • Training Efficiency: Achieved a 2x speedup during the fine-tuning process compared to standard methods.

Primary Focus

This model is a fine-tuned version of a model specifically designed for "answer-last" tasks, suggesting an optimization for generating responses or answers, potentially in a conversational or question-answering context. The fine-tuning process with Unsloth indicates an emphasis on efficient resource utilization during training.