jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random
The jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random model is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It features a substantial 32768-token context length, making it suitable for processing longer sequences of text. This model is a specialized iteration, having undergone further fine-tuning on an unspecified dataset, suggesting potential for niche applications beyond its base instruction-following capabilities.
Loading preview...
Model Overview
This model, jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random, is a fine-tuned variant of the Qwen2.5-1.5B-Instruct base model, developed by Qwen. It features 1.5 billion parameters and supports a 32768-token context length, enabling it to handle extensive textual inputs and generate coherent, long-form responses.
Key Characteristics
- Base Model: Built upon the robust Qwen2.5-1.5B-Instruct architecture.
- Fine-tuning: The model has undergone additional fine-tuning on an undisclosed dataset, indicating a specialization or adaptation for particular tasks or data distributions.
- Context Window: Its large context window is beneficial for tasks requiring deep understanding of lengthy documents or conversations.
Training Details
The fine-tuning process utilized specific hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation: 16 steps, resulting in a total effective batch size of 16.
- Optimizer: AdamW with default betas and epsilon.
- Scheduler: Linear learning rate scheduler with 200 warmup steps.
- Epochs: Trained for 3 epochs.
Intended Use
While specific intended uses and limitations are not detailed in the provided information, its instruction-tuned nature combined with further fine-tuning suggests potential for specialized instruction-following tasks where the additional training data has provided a performance boost or adapted it to a particular domain.