roskosmos19/Orca-2-4B-hight

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The roskosmos19/Orca-2-4B-hight is a 4 billion parameter Qwen3ForCausalLM-based model, optimized for agentic and coding tasks. It features a 32,768 token context window and is designed for faster, cheaper inference compared to its predecessors. This model excels in agentic workflows by utilizing a single-pass reasoning style with optional tokens and a clean tool-calling format.

Loading preview...

Orca-2-4B-hight: Optimized for Agentic & Coding Tasks

This model, based on the Qwen3ForCausalLM architecture with 4 billion parameters, is an optimized successor in the Athenea/Rhea coding lineage. It is specifically engineered for agentic and coding tasks, prioritizing faster and cheaper inference while maintaining strong performance.

Key Differentiators & Capabilities

  • Efficient Reasoning: Employs a single-pass reasoning style with optional <think>...</think> tokens, allowing the model to decide when chain-of-thought is beneficial, unlike previous multi-pass approaches.
  • Optimized for Agents: Features an improved tool-calling template and a clean, lean set of special tokens, making it highly suitable for agentic workflows.
  • Cost-Effective: Significantly reduces inference cost and latency by removing forced long outputs and artificial multi-pass overhead.
  • Strong Coding Focus: Retains a strong emphasis on coding and reasoning, encouraging precise, secure, and efficient solutions through its system prompt.
  • Generous Context Window: Offers a 32,768 token context length, ample for real-world agentic workloads.

When to Use This Model

This model is ideal for developers building applications that require:

  • Agentic AI systems needing reliable tool-calling and efficient reasoning.
  • Code generation and problem-solving where speed and cost are critical.
  • Applications benefiting from a large context window without the overhead of forced long generations.

It is recommended to use settings like temperature: 0.4 and top_p: 0.9 for optimal quality and speed, and quantizations like Q4_K_M or AWQ for best performance.