minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8
The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8 is a 4.5 billion parameter Qwen3.5-based language model, fine-tuned by minsu0567. This model was trained using Unsloth and Huggingface's TRL library, achieving a 2x speedup during the fine-tuning process. It is specifically fine-tuned from minsu0567/Uni-IAD-R2-Qwen3.5-answer-last, indicating a specialization in answer generation tasks.
Loading preview...
Model Overview
The minsu0567/Uni-IAD-R2-Qwen3.5-GRPO-answer_last-no-hard_total_batch_8 is a 4.5 billion parameter language model developed by minsu0567. It is based on the Qwen3.5 architecture and has been fine-tuned from the minsu0567/Uni-IAD-R2-Qwen3.5-answer-last model.
Key Characteristics
- Architecture: Qwen3.5 base model.
- Parameter Count: 4.5 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Fine-tuning Method: Utilizes Unsloth and Huggingface's TRL library for efficient training.
- Training Efficiency: Achieved a 2x speedup during the fine-tuning process compared to standard methods.
Primary Focus
This model is a fine-tuned version of a model specifically designed for "answer-last" tasks, suggesting an optimization for generating responses or answers, potentially in a conversational or question-answering context. The fine-tuning process with Unsloth indicates an emphasis on efficient resource utilization during training.