armand0e/Qwen3.5-9B-Fable-5-SDFT

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 17, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

The armand0e/Qwen3.5-9B-Fable-5-SDFT is a 9 billion parameter Qwen3.5 model fine-tuned using Self-Distillation Fine-Tuning (SDFT) on agentic coding and tool-use traces from Claude Fable 5. This model is specifically optimized for generating long-context assistant responses for coding agents and tool-use scenarios, supporting a context length of up to 32768 tokens. Its unique SDFT training method, which involves on-policy distillation, aims to improve performance on complex, multi-turn interactions.

Loading preview...

Model Overview

The armand0e/Qwen3.5-9B-Fable-5-SDFT is a 9 billion parameter language model built upon the unsloth/Qwen3.5-9B base. It distinguishes itself through its training methodology: Self-Distillation Fine-Tuning (SDFT). This technique uses the model itself in both student and teacher roles, where the student learns from the teacher's distribution, which is conditioned on expert demonstrations. Unlike traditional Supervised Fine-Tuning (SFT) that trains on fixed expert-written tokens (off-policy), SDFT trains on the model's own sampled tokens (on-policy), allowing for more dynamic and relevant learning.

Key Capabilities & Features

  • Agentic Coding & Tool-Use: Specifically fine-tuned on traces from Claude Fable 5, making it adept at generating responses for agentic coding tasks and tool utilization.
  • Long Context Handling: Supports a substantial context length of 32,768 tokens, with a prompt cap of 57,344 tokens and a rollout cap of 8,192 new tokens, suitable for complex, multi-turn conversations.
  • On-Policy Learning: Leverages SDFT to minimize divergence between the student's distribution and the expert-conditioned teacher's distribution, potentially leading to more robust and contextually aware outputs for interactive tasks.

When to Use This Model

  • Developing AI Agents: Ideal for applications requiring an AI assistant to perform agentic coding or interact with tools, where multi-turn reasoning and accurate response generation are critical.
  • Complex Conversational AI: Suitable for scenarios demanding long-context understanding and generation, such as detailed problem-solving or interactive development environments.
  • Research into Distillation Methods: Offers a practical implementation of the SDFT training method for those interested in exploring advanced fine-tuning techniques.

Limitations

It's important to note that the model's performance is influenced by the recorded traces it was trained on, potentially inheriting their errors or stylistic biases. While SDFT is on-policy per assistant turn, the overall environment feedback is still based on recorded expert trajectories, not live interaction. Tool calls generated by the model require downstream validation for safety and correctness. This model is not safety-tuned or policy-aligned for high-stakes decisions.