Dingdust/Ornith-1.5-9B-heretic

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 24, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Ornith-1.5-9B-heretic is a 9 billion parameter decensored version of the Ornith-1.5 model, created using the Heretic v1.4.0 framework. Developed by Ornith Team, this model is built on Qwen3.5 and Gemma4, featuring an end-to-end self-improvement loop for task generation and solution optimization. It excels in coding and reasoning tasks, demonstrating improved performance and reduced refusals compared to its original counterpart, and supports a 32768 token context length.

Loading preview...

Ornith-1.5-9B-heretic: Decensored, Self-Improving 9B LLM

This model is a 9 billion parameter variant of the Ornith-1.5 series, enhanced with the Heretic v1.4.0 framework to reduce content refusals. Developed by the Ornith Team, Ornith-1.5 represents a significant advancement in self-improvement, moving beyond fixed human-curated tasks to continuously generate new training tasks and discover effective solution strategies through reinforcement learning. It builds upon foundational models like Qwen3.5 and Gemma4 with extensive pre-training, mid-training, and post-training.

Key Capabilities & Differentiators

  • Decensored Output: Achieves a refusal rate of 33/100 compared to 84/100 in the original model, making it suitable for use cases requiring less restrictive content generation.
  • Enhanced Coding Performance: Demonstrates strong performance across various coding benchmarks, including Terminal-Bench 2.1 (46.2-47%), SWE-bench Verified (70.6%), SWE-bench Pro (47.5%), and SWE-bench Multilingual (54.4%), often outperforming its predecessor and other 9B models.
  • Advanced Reasoning: Shows improved reasoning capabilities on benchmarks like HLE (20.2% without tools, 30.5% with tools) and GPQA Diamond (86.4%).
  • Agentic Functionality: Optimized for agentic workflows, supporting tool calling and integration with frameworks like Ollama, Atomic.chat, llama.cpp, Hermes Agent, and OpenClaw.
  • Long Context Window: Supports a native context window of 262,144 tokens, extendable to approximately 1 million tokens using YaRN RoPE scaling for demanding long-context tasks.
  • Efficient Deployment: Designed as a dense 9B model for efficient single-GPU deployment, with a quantized mobile variant available.

Good For

  • Applications requiring less filtered or 'decensored' language generation.
  • Code generation, debugging, and understanding, especially for terminal-based coding agents.
  • Complex reasoning tasks and problem-solving.
  • Developing and deploying AI agents that utilize tool-use and long context windows.
  • Edge deployment on mobile devices via its quantized version.