holi-lab/ArcANE-32B-DPO

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ArcANE-32B-DPO is a 32 billion parameter Qwen3-based model developed by holi-lab, fine-tuned using supervised fine-tuning (SFT) and Direct Preference Optimization (DPO). This model is specifically designed for advanced character role-playing, excelling at maintaining character consistency and distinguishing subtle behavioral changes across narrative phases. It is optimized for point-in-time character responses conditioned on chapter-truncated character arcs, ensuring fidelity to the current narrative context.

Loading preview...

ArcANE-32B-DPO: Advanced Character Role-Playing Model

ArcANE-32B-DPO is a 32 billion parameter model built upon the Qwen/Qwen3-32B architecture, developed by holi-lab. It has undergone a two-stage training process: supervised fine-tuning (SFT) followed by Direct Preference Optimization (DPO). This unique training approach enables the model to learn to differentiate between a correct, in-character response at a specific narrative phase and a plausible but incorrect response from an adjacent phase.

Key Capabilities

  • Point-in-time Character Role-Play: Excels at generating responses that are consistent with a character's state at a precise moment in a narrative.
  • Contextual Character Responses: Conditions character responses on a chapter-truncated Character Arc, ensuring fidelity to the current story phase.
  • Subtle Behavioral Distinction: Capable of discerning and reproducing subtle behavioral changes between adjacent narrative phases, crucial for dynamic character development.
  • Improved Phase-Fidelity: DPO training significantly enhances the model's ability to maintain Action Phase-Fidelity (APF), Reasoning Phase-Fidelity (RPF), Reasoning-Action Entailment (RAE), and Phase Trajectory Fidelity (PTF) compared to its SFT-only counterpart and the base Qwen3-32B model.

Intended Use Cases

  • Research on Character Arc Modeling: Ideal for academic and research purposes focused on understanding and generating character behavior over narrative arcs.
  • Interactive Storytelling and Game Development: Can be used to power highly consistent and context-aware non-player characters (NPCs) in interactive narratives.
  • Creative Writing Assistance: Assists writers in maintaining character consistency and development throughout complex stories.

For optimal performance, the model should be supplied with the relevant Character Arc truncated up to the queried chapter, without exposing future narrative phases. The model was accepted to the EMNLP 2026 Main Conference, with further details available in the associated paper.