richardyoung/mythos-9b-unhinged-heretic

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The richardyoung/mythos-9b-unhinged-heretic is an 8 billion parameter decensored language model, derived from King3Djbl/mythos-9b-unhinged using the Heretic v1.4.0 framework. This model is a SLERP merge of Mythos-9B with braindao/Qwen3-8B-Uncensored, featuring a 32,768 token context length. It is specifically designed to provide complete answers without refusal across all topics, retaining core agent tool-use and reasoning capabilities. This model is optimized for use cases requiring maximum uncensored output and comprehensive responses.

Loading preview...

Overview

This model, richardyoung/mythos-9b-unhinged-heretic, is an 8 billion parameter decensored variant of the King3Djbl/mythos-9b-unhinged model, created using the Heretic v1.4.0 framework. It is a SLERP merge of Mythos-9B with braindao/Qwen3-8B-Uncensored at a 65% merge weight towards the uncensored partner. This configuration aims to remove almost all safety filters while preserving the original model's agentic capabilities.

Key Capabilities

  • Maximum Uncensorship: Designed to provide complete answers on every topic without refusal, achieving a 4.8/5 censorship resistance score.
  • Agentic Functionality: Retains core tool-use and reasoning capabilities from its base model.
  • Extended Context: Features a context length of 32,768 tokens.
  • Reproducible: The model's creation process is reproducible, with details available in the reproduce directory.
  • Performance: Offers approximately 8-10 tokens/second inference speed on an M3 Mac in --no-think mode.

Good For

  • Use cases requiring a model that never refuses to answer, regardless of the query's nature.
  • Applications where comprehensive, unfiltered responses are critical.
  • Developers needing an agent model with preserved tool-use and reasoning but without inherent safety filters.

Usage Notes

For faster responses, it is recommended to use the --no-think or /no_think mode, as the native thinking tokens consume output budget. The model is compatible with various platforms including Ollama, HuggingFace Transformers, vLLM, LM Studio, and llama.cpp.