richardyoung/Mythos-nano-heretic

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The richardyoung/Mythos-nano-heretic is a 3.1 billion parameter decensored version of the Mythos-nano model, created by Richard Young using Heretic v1.4.0. This model is specifically optimized for competitive programming and mathematical reasoning tasks, demonstrating strong performance on benchmarks like AIME and LeetCode. It features a 32768 token context length and has significantly reduced refusal rates compared to its original counterpart, making it suitable for use cases requiring less restrictive content generation.

Loading preview...

Model Overview

richardyoung/Mythos-nano-heretic is a 3.1 billion parameter language model, a decensored variant of squ11z1/Mythos-nano, developed by Richard Young using the Heretic v1.4.0 tool. This version has been modified to remove the refusal direction, resulting in a significantly lower refusal rate (13/100 compared to 79/100 for the original model) and reduced safety guardrails. It maintains a substantial context length of 32768 tokens.

Key Capabilities & Performance

This model excels in complex reasoning tasks, particularly in mathematics and competitive programming. Benchmarks highlight its strong performance:

  • Mathematics: Achieves scores like 91.4 on AIME25 and 94.3 on AIME26, with a variant (Mythos-nano + CLR) reaching 96.7 and 97.1 respectively. These scores are competitive with much larger, trillion-parameter systems.
  • Coding: Demonstrates a 96.1% pass-rate on LeetCode contests (Python), placing it among top-tier models like Gemini 3.1 Pro and GPT-5.2.
  • Decensored Output: The removal of the refusal direction allows for less restricted content generation, which users should handle responsibly.

Use Cases

  • Competitive Programming: Highly recommended for LeetCode-style problems and other programming challenges.
  • Mathematical Problem Solving: Ideal for tasks requiring advanced mathematical reasoning.
  • Unfiltered Content Generation: Suitable for applications where a model with reduced safety guardrails and lower refusal rates is desired, with the understanding that users are responsible for outputs and legal compliance.

Note: This model is not recommended for tool-calling, API orchestration, or agent-based programming due to its training data. Recommended sampling parameters include a temperature of 0.6โ€“1.0 and up to 40960 output tokens for challenging problems.