Gray Swan announced that the IPI benchmark was shown up in Google's Gemini 3.8 Flash Cyber launch.
Public source
Publisher name
Public post
Our IPI benchmark just showed up in Google's Gemini 3.8 Flash Cyber launch. I was suprised a few weeks back to see how much Gemini jumped up the leaderboards on our IPI…
Company
Gray Swan
Empowering the world to use AI safely and securely
- Industry
- Computer and Network Security
- Location
- Pittsburgh, US
- Company size
- 51–200 employees
About Gray Swan
Gray Swan AI is an AI safety and security company. We develop tools that automatically assess the risks of AI models and provide best-in-class safety and security.
See moreLatest activity
Latest activity from Gray Swan
9 signals
Products & Services
Gray Swan helped ensure the release of GPT-6 Astra safely and securely.
Partnerships
Gray Swan is used by AI labs to find model vulnerabilities before releasing models, with partners Anthropic, OpenAI, and Google DeepMind.
Partnerships
Gray Swan is the preferred adversarial testing and prevention partner for Anthropic with Fable 5.1, OpenAI announces GPT-6 Astra, and Google releases Gemini 3.8 Flash.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Ridge Security Technology Inc.
Ridge Security Technology Inc. published the industry's first public benchmark comparing leading LLMs in autonomous penetration-testing workflows.
Research & Knowledge
SnowCrash Labs
SnowCrash Labs reports that Astra is the first model OpenAI has graded Critical for cyber capability under its own Preparedness Framework after finding unknown flaws in a hardened browser and chaining them into a sandbox escape.
Research & Knowledge
Cequence Security
Cequence Security published a research claim that a global financial services firm ran their first AI discovery scan expecting a manageable number of agents, finding 740% more MCP servers, 800% more AI agents, and 600% more LLM providers than anticipated.
Research & Knowledge
Grip Security
Grip Security published an article looking beyond prompt injection at the full agentic AI attack surface and controls security teams need to limit what happens when an agent gets it wrong.
Research & Knowledge
Simbian