longtermrisk/Qwen3-8B-old-bird-names-last-third-v2-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-old-bird-names-last-third-v2-sft is an 8 billion parameter Qwen3 model, developed by longtermrisk, with a 32768 token context length. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is optimized for specific tasks related to its fine-tuning, building upon the base capabilities of Qwen3.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-old-bird-names-last-third-v2-sft, is an 8 billion parameter variant of the Qwen3 architecture. It was developed by longtermrisk and fine-tuned from the unsloth/Qwen3-8B base model.

Key Characteristics

  • Architecture: Qwen3-8B, a causal language model.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Training Efficiency: The fine-tuning process leveraged Unsloth and Huggingface's TRL library, resulting in a 2x speed improvement during training.

Intended Use

This model is suitable for applications requiring a Qwen3-8B base model that has undergone specific fine-tuning, particularly benefiting from the accelerated training methodologies employed. Its substantial context length makes it capable of handling longer inputs and generating more coherent, extended outputs for tasks aligned with its fine-tuning objective.