Paragon conducted tool-calling evaluations earlier this year measuring description quality, which is a related lever but a different one.
Public source
Publisher name
Public post
When an agent fails in production, the failure is usually silent. A paper out this week measured exactly how silent, and the numbers are worse than I expected. The easy…
Company
Paragon
Ship every integration your customers need with Paragon's integration infrastructure for AI and B2B software products.
- Industry
- Embedded Software Products
- Location
- Los Angeles, US
- Company size
- 51–200 employees
About Paragon
Paragon is the leading integration infrastructure platform for AI. Companies like Zendesk, Postman, and Five9 use Paragon to power integrations for agent tool calling, RAG ingestion, and more. With 130+ pre-built connectors, fully-managed authentication, and embedded SDK / APIs designed for native product integrations, Paragon helps developers scale integrations 10x faster than building in-house.
See moreLatest activity
Latest activity from Paragon
8 signals
Research & Knowledge
Paragon surveyed 600 B2B SaaS leaders this year, finding that 81% are now shipping or building AI agents that act through integrations.
Research & Knowledge
Paragon published the full 2026 State of Agentic Integrations report, consisting of all 600 responses from B2B SaaS leaders.
Research & Knowledge
Paragon surveyed 600 B2B SaaS leaders and found that two-thirds have lost or delayed an enterprise deal over integration security.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
SHIP
SHIP built a SWE-in-a-team coding benchmark to evaluate engineering delivery loops by comparing the cost of resolving tickets using different models.
Research & Knowledge
Sonar
Sonar published research showing that Armature measured which tools Claude Code, Codex, and Cursor actually install across 16,893 coding sessions.
Research & Knowledge
Glean
Glean's evaluation team runs early model checkpoints through real Glean workloads to measure quality, grounding, tool use, long-running execution, token efficiency, and cost.
Research & Knowledge
PlayerZero
PlayerZero rebuilt its harness around decisions to achieve a 30.7% increase in token efficiency.
Research & Knowledge
Tessl