Delta-Vector/Austral-32B-GLM4-Winton

Hugging Face
TEXT GENERATIONPricing:Input $1.06 / Cached $0.053 / Output $2.6Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 17, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

Delta-Vector/Austral-32B-GLM4-Winton is a 32 billion parameter GLM-4-Tulu based language model developed by Delta-Vector. This model is a multi-stage finetune specifically optimized as a generalist roleplay and adventure model. It excels at pushing plot narratives forward and improving general writing quality, having undergone KTO enhancement and training on natural human writing datasets.

Loading preview...

Austral 32B GLM4 Winton: A Roleplay and Adventure Specialist

Delta-Vector/Austral-32B-GLM4-Winton is a 32 billion parameter model built upon the Delta-Vector/GLM-4-32B-Tulu-Instruct base. This model has undergone a multi-stage finetuning process, including Knowledge Transfer Optimization (KTO) and training on diverse datasets like light novels and natural human writing, to enhance its coherence and writing quality.

Key Capabilities

  • Generalist Roleplay and Adventure: Specifically designed and optimized for engaging in roleplay scenarios and driving adventure narratives.
  • Enhanced Writing Quality: Finetuned to improve general writing and storytelling, aiming to push plot development effectively.
  • Coherency and Cohesiveness: Utilizes KTO to address and improve model coherency issues, resulting in a more consistent output.
  • GLM-4-Tulu Based: Leverages the architecture of GLM-4-Tulu, providing a robust foundation for its specialized finetuning.

Training Details

The model was trained over 4 epochs for the base Supervised Fine-Tuning (SFT), followed by 1 epoch of KTO for coherency, and a final epoch on a 'Rep_Remover' dataset to eliminate common model 'slops'. The total training time was approximately 80 hours, utilizing 8 x A100 GPUs.

Quantized Formats

Quantized versions are available for various deployment needs:

  • GGUF: For use with LLama.cpp and its forks.
  • EXL3: For use with TabbyAPI.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p