Proxion published notes on testing the grader using seeded error injection to evaluate its detection rate for defect classes.
Public source
Publisher name
Public post
Notes on testing the grader. A published evaluation result carries an implicit claim: the grading apparatus would have caught the error if there had been one. That claim…
Company
Proxion
Institutional financial reasoning data for frontier AI
- Industry
- Data Infrastructure and Analytics
- Location
- New York, US
- Company size
- 1 employee
About Proxion
Proxion turns institutional financial judgment into structured training and evaluation data for AI systems. Most finance content explains the facts. It does not capture the judgment behind real decisions: how practitioners underwrite risk, pressure test assumptions, build valuation logic, and reason through complex financial situations. Proxion captures that judgment. We work with credentialed finance professionals across investment banking, private equity, and asset management to produce training datasets, evaluation benchmarks, reasoning traces, rubrics, and preference data for financial AI systems. Our work helps AI systems move beyond financial retrieval toward financial reasoning. We support frontier AI labs, fintech platforms, and financial institutions building reliable, finance-grade AI. Based in New York.
See moreLatest activity
Latest activity from Proxion
6 signals
Research & Knowledge
Proxion published results from Scale AI's Scale Labs testing its own verifiers, where 38 professional office tasks were tested and zero of the 38 verifiers failed the document.
Research & Knowledge
Proxion published domain splits for APEX Agents, which runs the same agent through the day to day work of a corporate lawyer, a management consultant and an investment banking analyst.
Products & Services
Proxion launched Claude Opus 5, which tops MortgageTax.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Checksum.ai
Checksum.ai published the mechanics of the AI testing pipeline, detailing what happens between a test failing in CI and a pull request landing in a repository.
Research & Knowledge
Jank.AI
Jank.AI reports that AI models are still finding only 60% of code issues, but getting better with model updates.
Research & Knowledge
Sauce Labs
Sauce Labs published an article breaking down the math regarding AI-linked failures costing $1.2 to $1.3 trillion and the speed mismatch between human-speed code review and AI-speed code.
Research & Knowledge
Snorkel AI
Snorkel AI published new research from Justin Bauer on what changed in Terminal-Bench 4.0.
Research & Knowledge
Foundational