cactopus/Omega-Sapphira-L3.3-70B-v1.3

Hugging Face
TEXT GENERATIONPricing:Input $2.88 / Output $2.88Concurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:llama3.3Architecture:Transformer0.0K Featherless Exclusive Warm

Omega-Sapphira-L3.3-70B-v1.3 by cactopus is a 70 billion parameter Llama 3.3-based merge model with an extended 128K token context length. This model uniquely blends the structural coherence of 'Omega' with the prose quality of 'Sapphira' by applying distinct, depth-varying merge curves to attention and feed-forward blocks. It is designed for consistent, high-quality text generation across diverse scenarios, excelling in maintaining narrative stability and producing refined prose.

Loading preview...

Model Overview

cactopus/Omega-Sapphira-L3.3-70B-v1.3 is a 70 billion parameter language model built on the Llama 3.3 architecture, featuring an impressive 128K token context window. This model is a sophisticated merge of ReadyArt/L3.3-The-Omega-Directive-70B-Unslop-v2.1 and BruhzWater/Sapphira-L3.3-70b-0.2, designed to combine the strengths of both parent models.

Unique Merging Strategy

What sets Omega-Sapphira apart is its dynamic merging approach. Instead of a fixed ratio, it applies separate, depth-varying curves for attention and feed-forward (MLP) blocks across its 80 layers. This means:

  • Sapphira (prose quality) is weighted more heavily in the middle layers (peaking at 66% of MLP weights at layer 48) where prose style is primarily shaped.
  • Omega (structural coherence) maintains a dominant role in attention layers (never less than 50% at any depth) to ensure strong instruction adherence, discourse coherence, and resistance to looping.

This intricate balance aims to deliver both robust structural integrity and superior writing style.

Key Capabilities & Behavior

  • Consistent Performance: The model is noted for its stability and consistency across long sessions and diverse scenarios, requiring minimal configuration adjustments.
  • High-Quality Prose: Leverages Sapphira's strengths for refined and nuanced text generation.
  • Structural Stability: Benefits from Omega's contribution to maintain narrative flow, follow instructions, and avoid repetitive outputs.
  • Extended Context: Inherits a 128K token context length, suitable for complex and lengthy interactions.

Usage Notes

  • Uses the Llama 3 chat template.
  • The tokenizer is from the Omega base; users should ensure add_bos_token: true is explicitly set if their loader doesn't prepend <|begin_of_text|> by default.
  • Recommended sampler settings are provided, with Mirostat mode 2 suggested for experimentation.

Intended Use

This model is unaligned and intended for adult fiction, capable of engaging with explicit and violent material without refusal. Users are advised to be 18 or older and are responsible for their generated content.