Stash evaluated the latest version of its product against memory benchmarks and found that the judge, control, and prompts differ across models.
Public source
Publisher name
Public post
It's absolutely crazy how meaningless the most popular memory benchmarks largely are. Read all with a heavy grain of salt. We've been evaluating the latest version of ou…
Company
Stash
Knowledge bases your agents love
- Location
- San Francisco, US
- Company size
- 1 employee
About Stash
The knowledge base your agents love to use
See moreLatest activity
Latest activity from Stash
3 signals
Strategy & Corporate Development
Stash announced that DoorDash acquired Metis, a company founded by Aryan Shah and Aayush Sheth who had spent a year on the question of what an agent should remember and put them in charge of AI.
Products & Services
Stash pivoted to build an AI slide maker that can handle agents on Henry's machine and talk to agents on my machine, becoming the product now.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
LlamaIndex
LlamaIndex benchmarked over 92 tools on ParseBench.
Research & Knowledge
Kachi AI
Kachi AI published findings from Stanford's SALT Lab analyzing real Claude conversations, offering a comparison with Anthropic's Economic Index from 18 months ago.
Research & Knowledge
Microsoft
Microsoft published PAST-Bench, a benchmark that runs seven base models and four agent frameworks through 204 episodes on vs. off retained memory and checks whether agent memory gains are real.
Research & Knowledge
EverMind
EverMind conducted an injection ablation on LongMemEval-S using core-only retrieval, finding that core-only averaged 1.95 episodes per query.
Research & Knowledge
SmartMemory