ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-en
The ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-en model is a 1.7 billion parameter language model, fine-tuned from the Qwen3-1.7B architecture by Qwen. This model is specifically optimized for reasoning and instruction-following tasks, building upon its base model's capabilities. It is designed for general-purpose applications requiring robust language understanding and generation.
Loading preview...
Model Overview
This model, ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-en, is a fine-tuned variant of the Qwen3-1.7B base model developed by Qwen. With 1.7 billion parameters and a context length of 32768 tokens, it is designed for efficient language processing.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3-1.7B.
- Parameter Count: 1.7 billion parameters.
- Context Length: Supports a substantial context window of 32768 tokens.
- Training Epochs: Trained for 3 epochs.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 5e-06
- Optimizer: ADAMW_TORCH with betas=(0.9, 0.95) and epsilon=1e-12.
- Batch Size: A total training batch size of 32 (train_batch_size: 4, gradient_accumulation_steps: 8).
- Scheduler: Cosine learning rate scheduler with a warmup ratio of 0.1.
Intended Uses
While specific intended uses are not detailed in the provided README, as a fine-tuned instruction model, it is generally suitable for:
- General-purpose text generation.
- Instruction-following tasks.
- Applications requiring a balance of performance and efficiency for its size class.