Braintrust evaluated 1,329 current events questions across 4 models and 14 conditions, comparing You.com, provider built-in search, and no search.
Public source
Publisher name
Public post
Imagine doing your job without ever looking anything up on the internet. That's an agent without web search. Giving agents the web changed what they can do, especially o…
Company
Braintrust
Active observability for agents in production.
- Industry
- Software Development
- Location
- San Francisco, US
- Company size
- 51–200 employees
About Braintrust
Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.
See moreLatest activity
Latest activity from Braintrust
8 signals
Products & Services
Braintrust announced that the MCP server now exposes write tools for coding agents to author prompts, scorers, classifiers, configure the Topics pipeline, build monitor views, create alerts and scheduled jobs, and run evals.
Products & Services
Braintrust launched Patterns and Debugger, a tool for recurring behavior analysis and root cause finding.
Products & Services
Braintrust launched Loop, a tool for investigating and acting on datasets, scorers, prompts, monitoring, and more.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
You.com
You.com conducted an independent evaluation by Braintrust to determine the search configuration moves agent accuracy about 40x more than model choice does, with 4 models and 1,329 questions.
Research & Knowledge
Modern AI Inc.
Modern AI Inc. published its Q3 2026 State of AI Search report measuring 110 B2B SaaS brands across ChatGPT, Claude, Gemini, and Perplexity, and ran the same 50 buyer questions through Google search.
Research & Knowledge
FlowHunt
FlowHunt ran 1,905 prompts through five AI engines to analyze AI search response data.
Research & Knowledge
Centaur.ai
Centaur.ai analyzed 26,780 quality-filtered preference judgments from 132 labelers who judged 4,272 distinct pairings of 2,848 model responses to 712 consumer-health prompts.
Research & Knowledge
Tessl