AI Safety Research Fund learned that Astra reduced its monitorability.
Public source
Publisher name
Public post
What’s most notable is Astra’s reduction in monitorability. This doesn’t bode well for the future. CoT was sometimes unfaithful, but still useful. Astra doesn’t need CoT…
Company
AI Safety Research Fund
Let's ensure the future of AI is safe, sane, and human-aligned. Help us fund it.
- Industry
- Non-profit Organizations
- Company size
- 1 employee
About AI Safety Research Fund
The AI Safety Research Fund exists to solve one of the most urgent challenges of our time: ensuring that rapidly advancing artificial intelligence benefits humanity rather than endangers it. Help fund critical AI safety research today. Whether you can contribute $50/month or $500,000, every donation makes a difference. Monthly donations are especially valuable as they provide steady funding for our operations. Donations are tax-deductible through our 501(c)(3) fiscal sponsor, Institute for Education, Research, and Scholarships (IFERS).
See moreDiscover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Safe AI Foundation USA
Safe AI Foundation USA published a September technical report titled AI Security vs. AI Safety: They are not the same, highlighting that AI security is not safety and that progress in AI Safety is slower than that of AI Security.
Research & Knowledge
The AI Whistleblower Initiative (AIWI)
The AI Whistleblower Initiative (AIWI) reported that someone with knowledge of OpenAI's newly released Astra model architecture could make it harder to monitor reasoning.
Research & Knowledge
Safe AI Forum
Safe AI Forum published a new report titled Technical Safeguards Against Extreme Misuse of AI mapping five safeguard practices across three levels for developers.
Research & Knowledge
Apollo Research
Apollo Research successfully red-teamed Anthropic’s auto-mode with multiple labs.
Research & Knowledge
AI Safety Asia