AuroraSystem/Ficus-0.1-Agent-12B
AuroraSystem/Ficus-0.1-Agent-12B is a 12 billion parameter open-weight language model from the Ficus family, developed by Aurora System. Built on google/gemma-4-12B-it, it is fine-tuned for cybersecurity tasks, bilingual RU/EN contexts, and agentic scenarios with tool calling, featuring a 32768 token context length. It excels in cybersecurity domains like vulnerability analysis and web security auditing, achieving 73.00% on SecEval and 97.00% on SecQA v2. This model is intended for research and educational purposes as a text assistant, particularly for cybersecurity specialists.
Loading preview...
Ficus 0.1 Agent 12B: A Cybersecurity-Focused LLM
Ficus 0.1 Agent 12B, developed by Aurora System, is a 12 billion parameter open-weight language model based on google/gemma-4-12B-it. It is specifically fine-tuned for cybersecurity tasks and bilingual (Russian/English) contexts, with an impressive 32768 token context length. While it includes experimental agentic capabilities and a built-in thinking mode, developers explicitly state that Ficus 0.1 is an early test model and not production-ready for agentic pipelines, as its tool-calling functionality is currently unreliable and fabricates results.
Key Capabilities
- Cybersecurity Expertise: Strong performance in vulnerability analysis, source code auditing, binary exploitation (pwn/rev), applied cryptography, CVE breakdowns, and penetration testing methodology.
- Web Security Auditing: Proficient in analyzing OWASP Top 10 vulnerabilities such as SQL injection, JWT authentication bypass, CSRF, SSRF, and IDOR/BOLA.
- Reverse Engineering: Capable of analyzing x86_64/ARM assembly listings, recovering C pseudocode, and analyzing custom virtual machines.
- Reasoning: Features a built-in thinking mode (
<|channel>thought) for complex reasoning, though caution is advised due to potential degenerate loops. - Multimodal Input (Inherited): Supports image, audio, and video input, inherited from its base model, though fine-tuning was text-only and multimodal behavior was not validated.
Performance Highlights
Ficus 0.1 demonstrates significant improvements in cybersecurity benchmarks compared to its base model:
- SecEval: 73.00% (vs. 62.80% for base Gemma 4 12B)
- SecQA v1: 100.00% (matches base)
- SecQA v2: 97.00% (matches base)
- CyberMetric-500: 92.20% (vs. 89.80% for base)
It also shows gains in general knowledge benchmarks like MMLU (+7.2 points) and MMLU-Pro (+23.2 points) when thinking mode is enabled. However, it exhibits regressions in code generation (LiveCodeBench pass@1: 34.8% vs. 54.5% for base) and multilingual performance (ruMMLU -4.4, MMMLU-ZH -7.2).
Should I use this for my use case?
Use Ficus 0.1 Agent 12B as a text assistant for:
- Cybersecurity research and education: Ideal for analyzing vulnerabilities, auditing code, and understanding penetration testing methodologies.
- CTF competitions: Can assist with tasks related to pwn/rev, applied cryptography, and general security problem-solving.
- Starting point for further fine-tuning: Its specialized dataset makes it a strong foundation for developing more advanced cybersecurity-focused models.
Do NOT use this model for:
- Production agentic pipelines: The model's agentic capabilities are experimental and known to fabricate tool calls and results. Connecting it to real systems, filesystems, or networks is strongly discouraged.
- Code generation: It shows a significant regression in code generation performance compared to its base model.
- General-purpose multilingual tasks: There is a noted drop in performance for Russian and Chinese language benchmarks.