ermiaazarkhalili/Qwen3-8B-SFT-Fable5

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The ermiaazarkhalili/Qwen3-8B-SFT-Fable5 is an 8.2 billion parameter language model, a LoRA fine-tune of the unsloth/Qwen3-8B base model. It was supervised fine-tuned using the private ermiaazarkhalili/Fable-5-Complete-2M-Clean dataset, leveraging Unsloth and TRL for efficient training. This model is designed for instruction-following tasks, inheriting the Qwen3ForCausalLM architecture and supporting a 32768 token context length. Its primary strength lies in its specialized fine-tuning for specific instruction-following behaviors derived from its unique training data.

Loading preview...

Model Overview

ermiaazarkhalili/Qwen3-8B-SFT-Fable5 is an 8.2 billion parameter language model, built upon the unsloth/Qwen3-8B base model. It has undergone LoRA (Low-Rank Adaptation) supervised fine-tuning using a private dataset, ermiaazarkhalili/Fable-5-Complete-2M-Clean, with the aid of Unsloth and TRL libraries. The model inherits the apache-2.0 license from its base.

Key Capabilities

  • Instruction Following: Specifically fine-tuned for instruction-following tasks based on its unique training data.
  • Efficient Fine-tuning: Utilizes LoRA with a rank of 16 and alpha of 16, trained with QLoRA (4-bit precision) for efficient adaptation.
  • Context Length: Supports a maximum sequence length of 4096 tokens during training.

Training Details

The model was trained for 2 epochs with a learning rate of 0.0002 and an effective batch size of 8. Training loss observations showed a reduction from 0.9521 to 0.7824 over 94,254 steps. It's important to note that no downstream benchmark evaluations have been conducted, so performance claims are based solely on training loss.

Limitations

  • No Benchmarks: Lacks formal benchmark evaluations; performance is inferred from training loss only.
  • Inherited Biases: Carries the biases, knowledge cutoff, and potential failure modes of the base Qwen3-8B model.
  • Specialized Fine-tuning: Behavior outside the distribution of the Fable-5-Complete-2M-Clean dataset is untested.
  • Merged Adapters: LoRA adapters are merged into the base weights, preventing detachment from this specific fine-tune.