AE Studio published a paper by Ethan Roland, lead author of the paper with Anthropic, explaining why dangerous knowledge stays in the model after training is finished.
Public source
Publisher name
Public post
When an AI model "unlearns" something dangerous, where does that knowledge go? It stays in the model. Unlearning teaches it to play dumb when it sees the exact phrasing…
Company
AE Studio
Frontier AI Alignment Research, AI strategy, Production AI System Development
- Industry
- Software Development
- Location
- Marina del Rey, US
- Company size
- 51–200 employees
About AE Studio
AE Studio does frontier AI alignment research, AI strategy, and production AI system development. Bootstrapped. ~150 people. No outside investors. Founded in 2016 in Los Angeles. The independence means we fund alignment research without anyone telling us to stop and we only take on work we actually believe in. Our alignment work includes collaborations with DARPA and Anthropic, publications at top AI conferences, and active research on gradient routing, self-other overlap, endogenous steering resistance, multi-agent systems and several other neglected approaches. We think AI is going to be the most transformative technology of our lifetimes, and we want to make sure it goes well. That means doing the research now, while we still have time to get it right. We bring the same rigor we apply to frontier AI research to helping companies effectively leverage AI. We start with the business problem, not the technology. Technology is moving at a pace that is faster than ever before and we help companies figure out what to do with that: how to restructure teams around AI, where humans add the most value, which problems are worth solving and in what order. We've generated $6M+ in extra weekly revenue for airline companies, developed state-of-the-art neural decoders for brain-computer interface systems, and enabled education companies to serve 10x more students with a 90% contractor cost reduction.
See moreLatest activity
Latest activity from AE Studio
6 signals
Research & Knowledge
AE Studio found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol and less likely to include incriminating information in its CoT.
Research & Knowledge
AE Studio's modular pre-training paper with Anthropic was accepted to ICML 2026, one of the top venues in AI research.
Presence & Recognition
AE Studio released a work co-authored by James Bowler alongside Elias Stengel-Eskin, Hale Sirin, Newton Sander, Carlos Bonetti, Sasha Boguraev, and Simon Kirby
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Anthropic
Anthropic released a paper titled Hacker-Opus on what happens after training an alignment team.
Research & Knowledge
Apollo Research
Apollo Research successfully red-teamed Anthropic’s auto-mode with multiple labs.
Research & Knowledge
Hugging Face
Hugging Face published a paper by Meta titled Research Preference Models (RPMs) that tackles how to instill research taste in agents as they conduct experiments.
Research & Knowledge
Joyful Agents
Joyful Agents published a paper titled Temporal Dynamics of Memory Poisoning in Web3-Style LLM Agents that constructs a schema-constrained dataset of 2,614 multi-step attack trajectories spanning four attack families, chain poisoning, policy rewriting, backdoor triggering, and slow drift executed over shared persistent memory.
Research & Knowledge
PyMC Labs