borekboissy/Millesime-2026-4b-phase1
Millesime-2026-4b-phase1 is a 4 billion parameter language model developed by borekboissy, fine-tuned from Qwen3-4B-Instruct-2507. This model was trained on the millesime_202608_sft dataset, achieving a validation loss of 0.9014. It is designed for general language understanding and generation tasks, building upon the capabilities of its Qwen3 base.
Loading preview...
Model Overview
Millesime-2026-4b-phase1 is a 4 billion parameter language model developed by borekboissy. It is a fine-tuned iteration of the Qwen3-4B-Instruct-2507 base model, specifically trained on the millesime_202608_sft dataset.
Training Details
The model underwent training for 2 epochs, utilizing a learning rate of 1e-05 and a total batch size of 32 across 4 devices. Key training hyperparameters included:
- Optimizer: ADAMW_TORCH_FUSED with betas=(0.9, 0.999)
- LR Scheduler: Cosine with 0.1 warmup steps
- Final Validation Loss: 0.9014
Performance
During training, the model demonstrated a progressive reduction in validation loss:
- Epoch 0.4228: Validation Loss 0.9928
- Epoch 0.8457: Validation Loss 0.9309
- Epoch 1.2681: Validation Loss 0.9176
- Epoch 1.6909: Validation Loss 0.9042
- Epoch 2.0: Validation Loss 0.9014
This model is suitable for applications requiring a compact yet capable language model, leveraging the foundational strengths of the Qwen3 architecture.