roskosmos19/Orca-4B-Instruct
The roskosmos19/Orca-4B-Instruct is a 4.0 billion parameter instruction-tuned causal language model based on the Qwen3-4B architecture, optimized for superior price/performance. It features a 16,384-token context window and reduced KV-cache VRAM usage by removing multimodal special tokens. This model is designed for efficient deployment and provides focused, high-quality answers at minimal operational cost.
Loading preview...
Model Overview
The roskosmos19/Orca-4B-Instruct model, also known as Dolphin-4B-Instruct-0409, is a 4.0 billion parameter instruction-tuned language model built upon the robust Qwen3-4B architecture. Its primary distinction lies in its aggressive efficiency tuning, aiming to deliver the best price/performance ratio among 4B instruct models.
Key Optimizations & Capabilities
This model achieves its efficiency through several key modifications compared to the original Qwen3-4B:
- Optimized Context Window: Features a 16,384-token native context, which is deemed sufficient for over 95% of real-world use cases, significantly reducing VRAM requirements.
- Reduced KV-cache VRAM: By removing multimodal special tokens, the model is leaner, leading to much lower KV-cache VRAM consumption.
- Tuned Generation Defaults: Incorporates specific generation settings (temperature 0.5, top_p 0.85, top_k 30, repetition_penalty 1.05) to produce more precise, less repetitive, and higher-quality answers.
- Stronger Instruction Following: Enhanced system prompt tuning for improved instruction adherence.
Technical Specifications
- Architecture: Qwen3ForCausalLM
- Parameters: 4.0B
- Context Length: 16,384 tokens
- Recommended Dtype:
bfloat16orfloat16 - Recommended Quantization: Q4_K_M / AWQ / GPTQ for optimal speed and quality.
Ideal Use Cases
This model is particularly well-suited for applications where cost-efficiency and focused, high-quality text generation are paramount, especially in scenarios where a 16k context window is sufficient. Its optimizations make it an excellent choice for deployment on resource-constrained environments or for large-scale inference where operational costs need to be minimized.