ferrazzipietro/Qwen3-8B-reas-int-065-only-loss-noprompt-3epochs-en

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ferrazzipietro/Qwen3-8B-reas-int-065-only-loss-noprompt-3epochs-en is an 8 billion parameter language model, fine-tuned from the Qwen/Qwen3-8B architecture. This model was trained for 3 epochs with a learning rate of 5e-06 and a context length of 32768 tokens. Its specific fine-tuning objective and dataset are not detailed, suggesting a specialized but undisclosed application.

Loading preview...

Model Overview

This model, ferrazzipietro/Qwen3-8B-reas-int-065-only-loss-noprompt-3epochs-en, is an 8 billion parameter language model derived from the Qwen/Qwen3-8B base architecture. It has undergone a fine-tuning process over 3 epochs, utilizing a cosine learning rate scheduler with a warmup ratio of 0.1.

Training Details

The fine-tuning was conducted with a learning rate of 5e-06, a train_batch_size of 1, and a gradient_accumulation_steps of 32, resulting in an effective total_train_batch_size of 64. The optimizer used was ADAMW_TORCH with specific beta and epsilon parameters. The training environment included Transformers 4.57.0 and Pytorch 2.14.0+cu130, distributed across 2 GPUs.

Key Characteristics

  • Base Model: Qwen/Qwen3-8B
  • Parameter Count: 8 billion
  • Training Epochs: 3
  • Learning Rate: 5e-06
  • Optimizer: ADAMW_TORCH

Limitations

The specific dataset used for fine-tuning and the model's intended uses or primary differentiators are not detailed in the provided information. Users should exercise caution and conduct further evaluation to determine its suitability for specific tasks, as its specialized training objective remains undisclosed.