jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_origin
This model is jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_origin, a 1.5 billion parameter instruction-tuned causal language model. It is a fine-tuned version of Qwen/Qwen2.5-1.5B-Instruct, designed for general instruction following tasks. With a 32768 token context length, it offers substantial capacity for processing longer inputs. The model's specific differentiators and primary use cases are not detailed in the provided information.
Loading preview...
Model Overview
This model, jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_origin, is a 1.5 billion parameter instruction-tuned causal language model. It is a fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model, indicating its primary purpose is to follow instructions effectively. The model supports a substantial context length of 32768 tokens, allowing it to process and generate longer sequences of text.
Training Details
The model was fine-tuned using the following hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation Steps: 16, resulting in a total effective batch size of 16
- Optimizer: AdamW with default betas and epsilon
- LR Scheduler: Linear warmup for 200 steps, followed by a linear decay
- Epochs: 3
The training utilized Transformers 4.50.0, Pytorch 2.6.0+cu124, Datasets 3.4.1, and Tokenizers 0.21.0.
Limitations
Specific details regarding the dataset used for fine-tuning, intended uses, and limitations are not provided in the available documentation. Users should exercise caution and conduct further evaluation to determine its suitability for specific applications.