genevaquant/Hermes-4-70B-u

TEXT GENERATIONConcurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:llama3Architecture:Transformer Featherless Exclusive Cold

Hermes 4 70B is a 70 billion parameter Llama-3.1-based reasoning model developed by Nous Research, featuring a 32768 token context length. It is distinguished by its hybrid reasoning mode, which includes explicit deliberation segments, and significant improvements in math, code, STEM, logic, and creative writing. This model is optimized for producing format-faithful outputs, including valid JSON, and offers enhanced steerability and reduced refusal rates.

Loading preview...

Hermes 4 70B: A Frontier Reasoning Model

Hermes 4 70B, developed by Nous Research, is a 70 billion parameter model built on Llama-3.1. It introduces a hybrid reasoning mode that allows the model to deliberate internally using <think>…</think> segments before generating a response, enhancing its problem-solving capabilities. The model's training involved a significantly expanded post-training corpus of approximately 5 million samples and 60 billion tokens, blending reasoning and non-reasoning data.

Key Capabilities

  • Advanced Reasoning: Demonstrates massive improvements across math, code, STEM, logic, and creative writing tasks.
  • Structured Outputs: Trained to produce valid JSON for given schemas and to repair malformed objects, making it highly suitable for structured data generation.
  • Enhanced Steerability: Offers extreme improvements in steerability, notably reducing refusal rates and aligning with user values, as evidenced by its performance on the RefusalBench benchmark.
  • Function Calling & Tool Use: Supports function/tool calls within a single assistant turn, integrating them after its internal reasoning process.

When to Use This Model

  • Complex Problem Solving: Ideal for applications requiring deep reasoning, logical deduction, and multi-step problem-solving.
  • Structured Data Generation: Excellent for tasks needing precise, format-faithful outputs like JSON generation or schema adherence.
  • Creative and Technical Content: Suitable for generating high-quality content in creative writing, coding, and STEM fields.
  • Applications Requiring Alignment: When a model needs to be highly steerable and conform to specific user values without unnecessary refusals.