Castellum.AI evaluates agent reliability in three ways using Arbiter proof-of-value (POV) tests on client alert data.
Public source
Publisher name
Public post
An AI agent can have a 99% accuracy rate and still behave inconsistently. The accuracy score captures performance under a defined set of conditions. It does not show whe…
Company
Castellum.AI
AI agents for L1 and L2 alert resolution
- Industry
- Software Development
- Location
- New York, US
- Company size
- 11–50 employees
About Castellum.AI
Castellum.AI helps AML teams eliminate alert overload. With AI agents that resolve alerts across customer onboarding, payments and investigation workflows, we enable teams to process investigations 6x faster while eliminating 94% of false positives. Our Arbiter agents are trained on each institution's own internal policies and procedures, ensuring every decision reflects that institution's specific risk appetite. There's no one-size-fits-all model: our agents work the way your compliance team works. Arbiter agents can be deployed as a modular AI layer within existing workflows, or as part of a consolidated Castellum.AI solution that includes global risk data — sanctions, PEPs, adverse media and more — with real-time screening and monitoring.
See moreLatest activity
Latest activity from Castellum.AI
2 signals
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Resolve AI
Resolve AI's memory system learns from every investigation, proactively maps customer environments, and turned onboarding from weeks of back and forth into hours
Research & Knowledge
AstroKube
AstroKube wrote a framework for distinguishing three adoption models for AI agents: Model 7 as a conversational interface on top of existing products, Model 8 as an infrastructure built for others to run agents, and Model 9 as autonomous agents for SRE and ops.
Research & Knowledge
Coder
Coder researchers took 471 ordinary shell commands and rewrote them as 2,826 agent skill files, running them past two enterprise coding agents across 5,629 attempts.
Research & Knowledge
Braintrust
Braintrust evaluated 1,329 current events questions across 4 models and 14 conditions, comparing You.com, provider built-in search, and no search.
Research & Knowledge
Lasso