r3lax/Ornith-1.5-9B-heretic

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The r3lax/Ornith-1.5-9B-heretic is a 9 billion parameter language model, a decensored version of ornith-ai/Ornith-1.5-9B created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation. Developed by r3lax, this model maintains strong coding, reasoning, and agentic capabilities while exhibiting significantly reduced refusal rates compared to its original counterpart. It is designed for efficient single-GPU deployment and excels in complex coding tasks and agentic workflows, supporting a 32K token context window, extendable to 1M with YaRN scaling.

Loading preview...

r3lax/Ornith-1.5-9B-heretic: A Decensored, High-Performance Reasoning Model

This model is a 9 billion parameter variant derived from ornith-ai/Ornith-1.5-9B, specifically modified using the Heretic v1.2.0 tool and the Arbitrary-Rank Ablation (ARA) method. The primary differentiator of this "heretic" version is its significantly reduced refusal rate (18/100 compared to 99/100 for the original model), while maintaining the strong performance characteristics of the Ornith-1.5 family.

Key Capabilities

  • Decensored Output: Achieves a substantially lower refusal rate, offering more direct responses.
  • Advanced Reasoning: Excels in complex reasoning tasks, including those requiring tool use, as evidenced by strong scores on HLE (with tools) and GPQA Diamond benchmarks.
  • Robust Coding: Demonstrates high performance across various coding benchmarks like SWE-bench (Verified, Pro, Multilingual) and Terminal-Bench 2.1, making it suitable for code generation and understanding.
  • Agentic Workflows: Designed to integrate seamlessly with agent frameworks, supporting OpenAI-compatible tool calling and chain-of-thought reasoning.
  • Long Context Handling: Supports a native context window of 32,768 tokens, extendable up to 1,000,000 tokens using YaRN scaling for demanding long-context applications.

Good for

  • Developers requiring less restrictive model behavior: Ideal for applications where the original model's refusal rate might be a hindrance.
  • Complex code generation and analysis: Its strong performance on coding benchmarks makes it suitable for programming assistants and automated coding tasks.
  • Building AI agents: The model's agentic capabilities and tool-calling support are well-suited for developing sophisticated AI agents.
  • Applications needing extended context: When dealing with very long documents or extensive conversational histories, the YaRN-extended context window is beneficial.