SuperXyrex/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SuperXyrex/qwen3-finetuned is a 0.8 billion parameter language model, fine-tuned from Qwen/Qwen3-0.6B. This model was trained on an unspecified dataset, achieving a validation loss of 2.0564. It is a general-purpose language model, with specific applications and limitations yet to be detailed by the developer.

Loading preview...

Model Overview

SuperXyrex/qwen3-finetuned is a language model based on the Qwen3-0.6B architecture, developed by SuperXyrex. This version has been fine-tuned from the original Qwen/Qwen3-0.6B model, featuring 0.8 billion parameters and a context length of 32768 tokens.

Training Details

The model was trained over 3 epochs using a learning rate of 2e-05 and a total batch size of 16 (train_batch_size: 2, gradient_accumulation_steps: 8). The optimizer used was ADAMW_TORCH_FUSED. During training, the validation loss decreased from 2.3622 in the first epoch to 2.0564 by the third epoch.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-0.6B.
  • Parameter Count: 0.8 billion parameters.
  • Context Window: Supports a context length of 32768 tokens.
  • Training Objective: Achieved a final validation loss of 2.0564.

Intended Use Cases

Specific intended uses and limitations for this fine-tuned model are not yet detailed in the provided information. Developers should consider its base architecture and training loss when evaluating its suitability for general language tasks.