Anthropic released a paper titled Hacker-Opus on what happens after training an alignment team.
Public source
Publisher name
Public post
When I talk to customers about security with AI, I always explain it in two areas. Security is what you place outside the model: sandboxing, identity, permissions, netwo…
Company
Anthropic
Anthropic is an AI safety and research company working to build reliable, interpretable, and steerable AI systems.
- Industry
- Research Services
- Company size
- 5,001–10,000 employees
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.
See moreLatest activity
Latest activity from Anthropic
105 signals
Presence & Recognition
Anthropic is walking through workloads best suited for Claude Platform on AWS, Claude on Amazon Bedrock, Claude Enterprise in AWS Marketplace, and Claude Desktop on Bedrock on September 16 at 10am PT.
Presence & Recognition
Anthropic is hosting the Nightshift dinner for AI-native engineers with Creandum and Passionfroot on 9 September in Berlin.
Presence & Recognition
Anthropic is hosting the Anthropic x Kibo Ventures Madrid Founder Dinner on 9 September in Madrid.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Apollo Research
Apollo Research successfully red-teamed Anthropic’s auto-mode with multiple labs.
Research & Knowledge
AE Studio