emulsazib/Qwen3-4B-Instruct-2507-GRPO-CTI

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The emulsazib/Qwen3-4B-Instruct-2507-GRPO-CTI model is a 4-billion-parameter Qwen3-based causal language model developed by Emul Sazib. It is fine-tuned using Group Relative Policy Optimization (GRPO) to specialize in cyber threat intelligence (CTI) tasks. This model excels at summarizing reports, reasoning over indicators of compromise, and producing structured answers for CTI analysis, with a context length of up to 262,144 tokens.

Loading preview...

CyberSentinel CTI: Specialized for Threat Intelligence

CyberSentinel CTI is a 4-billion-parameter language model developed by Emul Sazib, specifically fine-tuned for cyber threat intelligence (CTI) tasks. Built upon the Qwen3-4B-Instruct-2507 base model, it leverages Group Relative Policy Optimization (GRPO) with reward functions that shape both output format and answer correctness.

Key Capabilities

  • CTI Task Specialization: Optimized for defensive security and threat-intelligence workflows.
  • Report Summarization: Efficiently summarizes threat-intelligence reports and security advisories.
  • IOC and TTP Reasoning: Capable of reasoning over indicators of compromise (IOCs), Tactics, Techniques, and Procedures (TTPs), and adversary behavior.
  • Structured Output: Designed to produce structured answers for CTI analysis tasks.
  • High Context Length: Inherits a substantial context window of up to 262,144 tokens from its base model.
  • Efficient Training: Fine-tuned using Unsloth and TRL, with LoRA adapters merged back into the 16-bit base weights.

Intended Use Cases

  • Assisting SOC analysts, threat researchers, and blue teams with drafting and enrichment.
  • Supporting defensive security operations and educational purposes.

Limitations

  • No formal held-out evaluation is reported; performance should be validated on specific CTI benchmarks.
  • Can hallucinate security facts; outputs require human verification.
  • Primarily English-centric; performance in other languages is not characterized.