Microsoft evaluated LLMs for secret scanning offline and achieved a 95 percent reduction in false positives while keeping recall within the defined guardrail.
Public source
Publisher name
Public post
🎯 How #GitHub Evaluates #LLMs Before They Reach Production 🧭 GitHub shares the lessons learned evaluating LLMs for real-world secret scanning, where a model that perfo…
Company
Microsoft
- Industry
- Software Development
- Location
- Redmond, US
- Company size
- 10,001+ employees
About Microsoft
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting in differences. Because impact matters. Microsoft operates in 190 countries and is made up of approximately 228,000 passionate employees worldwide.
See moreLatest activity
Latest activity from Microsoft
1,613 signals
Research & Knowledge
Microsoft published its latest Responsible AI Transparency Report showing progress over the past year and re-engineering its Responsible AI Standard to adapt to agentic AI.
Products & Services
Microsoft released Power Platform Admin v1.4 MCP connectors that add five capacity tools built on the Allocations By Environment API published in July.
Products & Services
Microsoft Marketplace offers private offers for software vendors to create customized pricing and purchasing terms for their organization.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Sumo Logic
Sumo Logic tested fine-tuning an LLM to generate security detection rules on a real Sigma detection rule problem.
Research & Knowledge
Sonar
Sonar ran Anthropic's latest model through its LLM evaluation framework, benchmarking against Claude Opus 4.8, and found that bug density dropped 14%, vulnerability density dropped 20%, and blocker-level security issues fell 75% compared to Opus 4.8.
Research & Knowledge
Semgrep
Semgrep benchmarked OpenAI's new Astra model against 13 points higher than Luna, scoring 100% on ExploitBench.
Research & Knowledge
SecondStack
SecondStack published a blog post titled LLM guardrails that catch PII and secrets before egress, and why the detector is rarely the hard part.
Research & Knowledge
Contrast Security