GLOBAL GAUNTLET published a detailed tech report and released code, data, and benchmarks for BenchMIRT, which applied Item Response Theory to MMLU, GPQA, and HellaSwag across models from OpenAI, Anthropic, Google, and Meta.
Published
Signal category
Research & Knowledge
Quote
“Researchers built BenchMIRT, applying Item Response Theory to tests like MMLU, GPQA, and HellaSwag across models from OpenAI, Anthropic, Google, and Meta. They found many datasets overweight easy, redundant items, blur reasoning vs. recall, and show inflated scores from data leakage.”
— John S.|GLOBAL GAUNTLET team
Company
GLOBAL GAUNTLET
Complexity, simplified.
- Industry
- Technology, Information and Internet
- Location
- San Francisco, US
- Company size
- 2 employees
Global Gauntlet operates at the intersection of AI, M&A, and capital allocation. We support companies navigating high-stakes decisions—identifying where AI creates real enterprise value, structuring build vs buy vs partner strategies, and ensuring those decisions translate into execution. Work spans AI strategy, acquisitions, partnerships, and applied systems, approached as a unified problem rather than separate disciplines. The focus is not advisory in isolation, but decision-making that holds under real-world conditions—across technical feasibility, financial constraints, and operational complexity. Experience includes $4B+ in publicly announced transactions across organizations such as Google, Fitbit, Intuit, Philips Healthcare, and Thermo Fisher Scientific. This work has involved evaluating opportunities, structuring strategic initiatives, and supporting execution across cross-functional environments. Recent efforts are focused on AI-driven strategy, market intelligence, and system deployment—helping organizations move from experimentation to production, and from fragmented initiatives to coherent systems. Global Gauntlet engages selectively with companies operating at inflection points—where strategy, capital, and execution must align.
Founded 2022