OrcaCyber Zero 1.5 Cybersecurity Model Launches With 1M Context
OrcaRouter releases OrcaCyber Zero 1.5, a gated cybersecurity model with 1M context window, reporting 100% on Cybench and 95.8% on CVE-Bench.

What changed
OrcaRouter released OrcaCyber Zero 1.5 on October 10, 2026, a specialized cybersecurity model designed for vulnerability research, exploit development, and penetration testing. The model is the successor to OrcaCyber Zero 1.0, which shipped on September 17, 2026. According to the company, OrcaCyber Zero 1.5 is post-trained for security research and authorized security engineering.
The model ships with a 1M-token context window, native function calling, and structured outputs. It supports text input and text output, with a maximum output of 128K tokens. The company positions it as built for complex, multi-step security workflows with long-context reasoning for large codebases.
Access is gated to the Security Research tier through OrcaRouter. Trusted security researchers, red teams, and authorized security testing organizations must provide an engagement, a passkey, and accept the platform’s terms to gain access. The model is served through an OpenAI-compatible API.
Pricing is set at $3.00 per 1M input tokens and $7.50 per 1M output tokens. Cache reads cost $0.75 per 1M tokens. According to the model page, over the past 7 days, the p50 time-to-first-token was 3.16 seconds and p95 was 8.88 seconds, measured across 2.4B tokens of traffic.
Why it matters
Security teams now have a dedicated AI model explicitly trained for vulnerability analysis rather than relying on general-purpose language models. The company states that the model’s design goals are to find unknown flaws, reason through attack paths, and go beyond detection by challenging its own hypotheses and ranking flaws by demonstrable exploitability.
The 1M context window matters for security workflows because it allows the model to reason over entire codebases, dependencies, and related attack surface in a single context. For teams conducting penetration tests, security audits, or vulnerability reproduction, this means fewer context window limitations compared to models with smaller windows.
The vendor-reported 100% score on Cybench suggests strong performance on professional capture-the-flag tasks relevant to offensive security research. The 95.8% result on the CVE-Bench evaluable subset indicates capability on real-world critical-severity web vulnerabilities, though the subset tested was 24 of 40 total CVE-Bench tasks.
The gating mechanism carries significance. By restricting access to authorized security researchers and red teams, OrcaRouter positions this as a tool for sanctioned testing. This differs from open-access models and reflects the company’s apparent design philosophy around models with offensive capability.
Pricing positions the model at $3.00 per 1M input tokens and $7.50 per 1M output tokens. The company notes that this is far below Mythos Preview’s reported rates of $25 per 1M input tokens and $125 per 1M output tokens.
What to test
Before adopting OrcaCyber Zero 1.5, security teams should verify these vendor claims independently:
Benchmark validity: All reported scores are self-reported by OrcaRouter. The CVE-Bench result covers only 24 of 40 tasks, not the full benchmark. Independently replicate testing on representative vulnerability datasets to confirm performance on your organization’s vulnerability types.
Latency in production: The 3.16-second p50 time-to-first-token and 8.88-second p95 are measured over 7 days with 2.4B tokens of traffic. Test with your actual traffic patterns and workload distribution. Determine whether latency meets your incident response timing requirements.
Consistency on false positives: Test how often the model flags non-exploitable issues as critical. The company claims the model ranks flaws by demonstrable exploitability, but verify this does not generate excessive false positives that waste analyst time.
Context window in practice: While the model supports 1M tokens, test real-world codebases and dependency trees to confirm the context window handles your typical vulnerability assessment scope. Measure token consumption on representative security audits.
Access gate enforcement: Confirm that access controls work as stated. Test that the passkey and engagement requirements are enforced consistently and that rate limits align with your OrcaRouter plan tier.
Comparison with general-purpose models: Benchmark OrcaCyber Zero 1.5 against your current tools on a sample of your vulnerability backlog to measure improvement in speed, accuracy, or reasoning quality.
SWE-bench gap: The model’s 76.5% on SWE-bench Pro V2 is notably lower than its cyber benchmark scores. Understand whether this general code understanding limitation affects exploit development or remediation code generation tasks relevant to your team.
The conclusion
OrcaCyber Zero 1.5 represents a narrowing of AI capability toward a specific professional domain. The 1M context window and native function calling address real constraints in vulnerability research workflows. The reported benchmark performance is strong, but all results are vendor-reported with no independent technical report yet published.
The gating mechanism and pricing structure suggest OrcaRouter is building a paid service for authorized security work rather than releasing an open tool. This is a deliberate tradeoff between accessibility and governance. Security teams serious about specialized AI for vulnerability analysis should test it within the access tier, benchmark it against their current tools, and verify performance on their own vulnerability types before committing to production use.
Watch for: independent replications of the Cybench and CVE-Bench results, publication of a technical report detailing training methodology and safety considerations, and whether the model’s reasoning quality sustains under adversarial testing by determined attackers.
AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure
Related stories

Odyssey-3 World Model Public Preview Launches Oct 11
Odyssey AI launched a public research preview of Odyssey-3, a 14B generative world model that creates interactive environments from text prompts in real time.

Claude Dynamic Workflows: 1,000 Parallel Agents Now Available
Anthropic added dynamic workflows to Claude Managed Agents, enabling up to 1,000 AI agents to run in parallel per execution through managed agent infrastructure.

Liquid AI Releases Multimodal Open D1 Decision Models for Edge
Liquid AI released d1-3B and d1-omni-600M, open multimodal decision models designed for edge devices, supporting text, images, and audio inputs.