borekboissy/Millesime-2026-4b-phase2
Millesime-2026-4b-phase2 is a 4 billion parameter language model developed by borekboissy, fine-tuned from Qwen/Qwen3-4B-Instruct-2507. This model was fine-tuned using the comparia_dpo dataset, achieving a reward accuracy of 0.7660. It is designed for tasks benefiting from instruction-tuned models, with a context length of 32768 tokens.
Loading preview...
Overview
Millesime-2026-4b-phase2 is a 4 billion parameter language model developed by borekboissy. It is a fine-tuned version of the Qwen/Qwen3-4B-Instruct-2507 base model, specifically optimized using the comparia_dpo dataset. This fine-tuning process aimed to enhance its performance, as indicated by the evaluation results.
Key Capabilities & Performance
The model demonstrates specific performance metrics on its evaluation set:
- Reward Accuracy: Achieved 0.7660.
- Reward Margin: Recorded at 4.2381.
- Loss: Final training loss was 0.8321.
These metrics suggest its proficiency in tasks aligned with its DPO (Direct Preference Optimization) fine-tuning, where it learns from preferred and rejected responses. The model was trained with a learning rate of 5e-06 over 1.0 epoch, utilizing a total batch size of 32 across 4 GPUs.
Training Details
The training procedure involved specific hyperparameters:
- Optimizer: ADAMW_TORCH_FUSED
- Learning Rate Scheduler: Cosine type with 0.1 warmup steps.
- Frameworks: Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, Tokenizers 0.22.2.
Intended Uses
While specific intended uses and limitations require further information, its fine-tuning on a preference dataset suggests suitability for tasks requiring nuanced understanding of preferred responses, such as instruction following, dialogue generation, or content moderation where distinguishing between good and bad outputs is crucial.