minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary
minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary is a 4.5 billion parameter Qwen3.5-based language model developed by minsu0567. This model was fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last and optimized for training speed using Unsloth and Huggingface's TRL library. It is designed for tasks related to answering the last part of a query, leveraging its 32768 token context length.
Loading preview...
Overview
minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard-type-binary is a 4.5 billion parameter language model built upon the Qwen3.5 architecture. Developed by minsu0567, this model is a fine-tuned iteration of the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last base model.
Key Characteristics
- Base Model: Qwen3.5 architecture.
- Parameter Count: 4.5 billion parameters.
- Training Optimization: The model's training process was significantly accelerated, achieving 2x faster training speeds by utilizing the Unsloth library in conjunction with Huggingface's TRL library.
- Context Length: Features a substantial context window of 32768 tokens.
Primary Use Case
This model is specifically fine-tuned for tasks that involve extracting or generating the "answer last" component of a given input. Its optimization for training speed and substantial context length make it suitable for applications requiring efficient fine-tuning and processing of longer sequences to identify final answers.