borekboissy/Millesime-2026-4b-phase1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

Millesime-2026-4b-phase1 is a 4 billion parameter language model developed by borekboissy, fine-tuned from Qwen3-4B-Instruct-2507. This model was trained on the millesime_202608_sft dataset, achieving a validation loss of 0.9014. It is designed for general language understanding and generation tasks, building upon the capabilities of its Qwen3 base.

Loading preview...

Model Overview

Millesime-2026-4b-phase1 is a 4 billion parameter language model developed by borekboissy. It is a fine-tuned iteration of the Qwen3-4B-Instruct-2507 base model, specifically trained on the millesime_202608_sft dataset.

Training Details

The model underwent training for 2 epochs, utilizing a learning rate of 1e-05 and a total batch size of 32 across 4 devices. Key training hyperparameters included:

  • Optimizer: ADAMW_TORCH_FUSED with betas=(0.9, 0.999)
  • LR Scheduler: Cosine with 0.1 warmup steps
  • Final Validation Loss: 0.9014

Performance

During training, the model demonstrated a progressive reduction in validation loss:

  • Epoch 0.4228: Validation Loss 0.9928
  • Epoch 0.8457: Validation Loss 0.9309
  • Epoch 1.2681: Validation Loss 0.9176
  • Epoch 1.6909: Validation Loss 0.9042
  • Epoch 2.0: Validation Loss 0.9014

This model is suitable for applications requiring a compact yet capable language model, leveraging the foundational strengths of the Qwen3 architecture.