holi-lab/ArcANE-8B-DPO

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ArcANE-8B-DPO by holi-lab is an 8 billion parameter Qwen3-8B model fine-tuned with supervised fine-tuning (SFT) and Direct Preference Optimization (DPO). It specializes in point-in-time character role-play, distinguishing subtle behavioral changes between adjacent narrative phases within a character arc. The model is designed for research into character responses conditioned on chapter-truncated character arcs, ensuring in-character consistency over time.

Loading preview...

ArcANE-8B-DPO: Character Arc-Aware Role-Playing Model

ArcANE-8B-DPO is an 8 billion parameter model developed by holi-lab, based on the Qwen3-8B architecture. It has been specifically trained using a two-stage process: Supervised Fine-Tuning (SFT) followed by Direct Preference Optimization (DPO). This training methodology enables the model to accurately distinguish between correct character responses at a specific narrative phase and plausible but incorrect responses from adjacent phases.

Key Capabilities

  • Point-in-time character role-play: Generates responses that are consistent with a character's development at a precise moment in a narrative.
  • Character Arc conditioning: Conditions responses on a chapter-truncated Character Arc, ensuring fidelity to the character's evolving personality and motivations.
  • Phase distinction: Excels at identifying and generating responses that reflect subtle behavioral changes between different narrative phases.
  • Research focus: Intended for research into advanced role-playing language agents and narrative consistency.

Training and Performance

The model was fine-tuned on the ArcANE corpus, utilizing 12 training novels, 55 characters, and 339 character axes. The DPO stage involved 14,671 preference pairs, where chosen responses belonged to the anchor phase and rejected responses to an adjacent phase. Evaluation showed ArcANE-8B-DPO with Arc context achieved an Overall score of 56.9, outperforming ArcANE-8B-SFT (52.3) and the base Qwen3-8B (43.1) under the same conditions. The model's training involved full fine-tuning with a maximum sequence length of 8,192 tokens for both SFT and DPO stages.

Recommended Use

This model is best utilized for research applications requiring highly nuanced, phase-aware character role-play. It is crucial to supply only the relevant Character Arc up to the queried chapter, ensuring future narrative phases are not exposed to the model for faithful point-in-time conditioning. The model uses the Qwen3 chat template with thinking disabled.