Fugu-Cyber
Sakana AI's cybersecurity-specialized orchestration model, claiming state-of-the-art scores on real-world security benchmarks like CyberGym and CTI-REALM.
Updated 2026-07-21
Overview
Fugu-Cyber is Sakana AI's orchestration model tuned specifically for cybersecurity work — vulnerability discovery, exploit reasoning, and threat-intelligence analysis — released July 21, 2026 as a new API endpoint. Where general frontier models handle security tasks as one capability among many, Fugu-Cyber is trained and evaluated against domain benchmarks, and Sakana positions its scores on CyberGym (real-world vulnerability tasks) and CTI-REALM (cyber threat intelligence) as state-of-the-art, matching the dedicated cyber variants of GPT-5.5.
It's built for security engineers, red teams, and threat researchers who want a model that behaves like a coordinated agent across security tooling rather than a chat window. As an "orchestration" model, the pitch is that it plans and sequences multi-step security workflows — enumerating a target, reasoning about a vulnerability class, drafting analysis — instead of answering one prompt at a time. This is a raw model accessed through an API, not a packaged app, so you're wiring it into your own pipeline or agent framework.
The honest caveat: nearly everything known about performance comes from Sakana's own release, and access is gated behind an application-and-approval process with no free tier or public playground. That makes independent verification hard right now, and the specialization cuts both ways — this is not a general-purpose model you'd reach for outside security contexts.
Key features
Cyber Orchestration
Positioned as an orchestration model that plans and sequences multi-step security workflows rather than answering isolated prompts, aimed at agentic red-team and analysis pipelines.
Frontier Benchmarks
Sakana reports state-of-the-art results on CyberGym (real-world vulnerability tasks) and CTI-REALM (threat intelligence), claiming parity with GPT-5.5's dedicated cyber variants.
API Endpoint
Delivered as a new API endpoint with token-based pricing, meant to be integrated into existing security tooling and agent frameworks rather than used through a chat UI.
Gated Access
Availability requires an application and approval, signaling a controlled rollout appropriate to a model with offensive-security capabilities.
Pricing
| Plan | Price | What's included |
|---|---|---|
| Token Plan | $6–$12/M input, $36–$54/M output | Pay-per-token API access; rates scale up for large-context requests. Access requires application approval. |
Pay-per-token API access; rates scale up for large-context requests. Access requires application approval.
Pros & cons
Pros
- ✓Specialized for cybersecurity rather than a general model used off-label
- ✓Claimed SOTA on real-world benchmarks (CyberGym, CTI-REALM), not just synthetic tests
- ✓Orchestration design targets multi-step agentic security workflows
- ✓Application-gated rollout is a reasonable posture for offensive-security capability
Cons
- ×Access is gated behind application approval with no free tier or public playground
- ×Performance claims come almost entirely from Sakana's own release — little independent verification yet
- ×Output pricing is steep and scales higher for large-context requests
- ×Narrow specialization: not useful outside security work
How it compares
| Tool | Best for | Pricing | Score |
|---|---|---|---|
| Fugu-Cyber | — | Token plan only: $6–$12/M input, $36–$54/M output (application required, no free tier) | 8.2/10 |
| Perplexity AI vs Perplexity AI → | — | Freemium | 9.4/10 |
| NotebookLM vs NotebookLM → | — | Free | 9.1/10 |
| Phind vs Phind → | — | Free tier + Pro subscription for advanced models | 8.7/10 |
Compare head-to-head
Related reading
Anthropic's $50K Claude Grants for Rare Disease Work
Anthropic is giving up to $50,000 in Claude API credits to rare-disease researchers, with applications closing August 2, 2026. Here's what's on offer.
OpenAI's Long-Horizon Safety Lessons, Unpacked
OpenAI's July 20 post on long-horizon model safety moves the unit of alignment from single actions to whole agent trajectories. Here's what it says.
GPT-Red vs Human Red Teaming: What Benchmarks Show
OpenAI's GPT-Red automates red teaming at scale, but human testers still find the attacks it can't. A neutral read of the numbers.
Ready to try Fugu-Cyber?
Head to the official site to start with Fugu-Cyber — pricing and plans are listed above.
Visit Fugu-Cyber