minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last_2
The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last_2 is a 4.5 billion parameter Qwen3.5 model developed by minsu0567. This model was fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last, leveraging Unsloth and Huggingface's TRL library for 2x faster training. It features a 32768 token context length and is optimized for specific answer generation tasks.
Loading preview...
Model Overview
The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last_2 is a 4.5 billion parameter language model, fine-tuned by minsu0567. It is based on the Qwen3.5 architecture and was specifically trained from the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last model.
Key Characteristics
- Architecture: Qwen3.5 base model.
- Parameter Count: 4.5 billion parameters.
- Context Length: Supports a substantial context window of 32768 tokens.
- Training Efficiency: The fine-tuning process was significantly accelerated, achieving 2x faster training speeds by utilizing Unsloth and Huggingface's TRL library.
Intended Use
This model is designed for applications requiring efficient and focused answer generation, likely benefiting from its specialized fine-tuning. Its optimized training process suggests a focus on performance and resource efficiency for its specific task.