XDQAQ/Qwen3-8B-MaKTO-Public-SFT

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

XDQAQ/Qwen3-8B-MaKTO-Public-SFT is an 8 billion parameter Qwen3-based model, fine-tuned using supervised learning on a public Chinese dataset focused on Werewolf game reasoning. This model specializes in generating responses for language-game and Werewolf-style interactions, offering a unique capability for strategic and context-aware dialogue within this specific domain. It is an independent reproduction inspired by the MaKTO paper, specifically adapted for Chinese Werewolf game scenarios.

Loading preview...

Model Overview

XDQAQ/Qwen3-8B-MaKTO-Public-SFT is an 8 billion parameter model based on the Qwen3 architecture. It has undergone full-parameter supervised fine-tuning (SFT) using a public Chinese dataset, specifically ReneeYe/werewolf_game_reasoning.

This model is an independent reproduction inspired by the MaKTO paper, focusing on the SFT stage. It's important to note that it does not include the complete MaKTO training recipe, as the general-purpose SFT examples and the Multi-agent KTO stage described in the paper were not utilized.

Key Capabilities

  • Specialized for Werewolf Game Reasoning: The model is fine-tuned on 12,886 examples, including game-behavior, advanced-technique, and fundamental-term examples from the Werewolf game domain.
  • Chinese Language Focus: Training data is exclusively from the public Chinese portion of the Werewolf game reasoning dataset.
  • Qwen3 Base: Leverages the capabilities of the Qwen3-8B base model.

Training Details

The model was trained using full-parameter BF16 SFT over 3 epochs, with a maximum sequence length of 6,144 tokens. It utilized 4 GPUs with DeepSpeed ZeRO-3 and a learning rate of 1e-6.

Limitations and Use Cases

This model is highly specialized for language-game and Werewolf-style interactions. It has not been evaluated as a general-purpose assistant and may produce incorrect, inconsistent, or strategically deceptive outputs, consistent with its training domain. Users should specifically consider this model for applications requiring nuanced, context-aware responses within Chinese Werewolf game simulations or similar strategic dialogue environments. It is not recommended for general conversational AI tasks.