Alu017/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Alu017/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from Qwen/Qwen3-0.6B. This model is a smaller variant within the Qwen3 family, featuring a 32768-token context length. It is designed for general language tasks, with its specific fine-tuning dataset and primary differentiators currently unspecified beyond its base architecture and parameter count.

Loading preview...

Model Overview

Alu017/qwen3-finetuned is a 0.8 billion parameter language model, derived from the Qwen/Qwen3-0.6B base model. It was fine-tuned over 2 epochs using a linear learning rate scheduler and AdamW_Torch_Fused optimizer, achieving a validation loss of 2.8675. The training utilized mixed-precision (Native AMP) with a total batch size of 16.

Key Training Details

  • Base Model: Qwen/Qwen3-0.6B
  • Parameters: 0.8 billion
  • Context Length: 32768 tokens
  • Learning Rate: 2e-05
  • Optimizer: AdamW_Torch_Fused
  • Epochs: 2
  • Final Validation Loss: 2.8675

Current Limitations

Specific details regarding the fine-tuning dataset, intended uses, and unique capabilities or differentiators of this particular fine-tuned version are not provided in the available documentation. Users should be aware that its primary strengths and ideal applications are currently undefined beyond its foundational Qwen3 architecture.