RISys-Lab/RedSage-K-GRPO

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026Architecture:Transformer Featherless Exclusive Cold

RedSage-K-GRPO is an 8 billion parameter cybersecurity model developed by RISys-Lab at Khalifa University, based on the Qwen3ForCausalLM architecture with a 32768 token context length. It specializes in translating natural-language requests into Kali/Linux commands, utilizing Group Relative Policy Optimization (GRPO) on the KaliBench dataset. This model is optimized for cybersecurity tool use, achieving an average Total Score of 76.9% on the KaliBench test split for command generation.

Loading preview...

Model Overview

RISys-Lab's RedSage-K-GRPO is an 8 billion parameter language model specifically designed for cybersecurity applications. Developed by Khalifa University, this model translates natural language queries into Kali/Linux commands, leveraging the Qwen3ForCausalLM architecture. It was trained using Group Relative Policy Optimization (GRPO) on the KaliBench dataset, a fine-grained benchmark for cybersecurity tool use.

Key Capabilities & Features

  • Cybersecurity Command Generation: Translates natural language requests into accurate Kali/Linux shell commands.
  • GRPO Training: Applies Group Relative Policy Optimization directly to its base model, RedSage-Qwen3-8B-Ins, using verifiable rewards on KaliBench without a preceding Supervised Fine-Tuning (SFT) stage.
  • Reasoning Output: Generates reasoning in <think> tags before producing the final command in an <output> tag.
  • KaliBench Performance: Achieves an average Total Score of 76.9% on the 5,000-example KaliBench test split, representing a +5.2 percentage point gain over its RedSage-Ins baseline. This includes strong performance in tool accuracy, optional-argument F1, and positional-argument F1.
  • Runtime-Free Verification: Rewards are computed from command structure without requiring execution of generated commands.

Use Cases & Considerations

RedSage-K-GRPO is intended for cybersecurity research, education, and command assistance within authorized environments. It excels at single-command generation for Kali/Linux tools. Users should be aware that generated commands require review before execution, as they may contain inaccuracies. The model's performance is measured on single-command generation and does not account for multi-step agent performance or execution success. It is not evaluated for multilingual, long-context, general-chat, or misuse-resistance capabilities.