roskosmos19/Orca-4B-Instruct

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The roskosmos19/Orca-4B-Instruct is a 4.0 billion parameter instruction-tuned causal language model based on the Qwen3-4B architecture, optimized for superior price/performance. It features a 16,384-token context window and reduced KV-cache VRAM usage by removing multimodal special tokens. This model is designed for efficient deployment and provides focused, high-quality answers at minimal operational cost.

Loading preview...

Model Overview

The roskosmos19/Orca-4B-Instruct model, also known as Dolphin-4B-Instruct-0409, is a 4.0 billion parameter instruction-tuned language model built upon the robust Qwen3-4B architecture. Its primary distinction lies in its aggressive efficiency tuning, aiming to deliver the best price/performance ratio among 4B instruct models.

Key Optimizations & Capabilities

This model achieves its efficiency through several key modifications compared to the original Qwen3-4B:

  • Optimized Context Window: Features a 16,384-token native context, which is deemed sufficient for over 95% of real-world use cases, significantly reducing VRAM requirements.
  • Reduced KV-cache VRAM: By removing multimodal special tokens, the model is leaner, leading to much lower KV-cache VRAM consumption.
  • Tuned Generation Defaults: Incorporates specific generation settings (temperature 0.5, top_p 0.85, top_k 30, repetition_penalty 1.05) to produce more precise, less repetitive, and higher-quality answers.
  • Stronger Instruction Following: Enhanced system prompt tuning for improved instruction adherence.

Technical Specifications

  • Architecture: Qwen3ForCausalLM
  • Parameters: 4.0B
  • Context Length: 16,384 tokens
  • Recommended Dtype: bfloat16 or float16
  • Recommended Quantization: Q4_K_M / AWQ / GPTQ for optimal speed and quality.

Ideal Use Cases

This model is particularly well-suited for applications where cost-efficiency and focused, high-quality text generation are paramount, especially in scenarios where a 16k context window is sufficient. Its optimizations make it an excellent choice for deployment on resource-constrained environments or for large-scale inference where operational costs need to be minimized.