We need more safety researchers
You're an engineer who always wanted to do a startup. You're searching for exciting problems to solve. I want to propose to you an alternative path to doing the next Facebook: AI safety.
This is the most important problem of our lifetime, and it's a great one to pick if you're young and ambitious.
Last century brought us the airplane, the internet, computers, the moon landing and the most innovation we've ever experienced, yet the most important man was Robert Oppenheimer. We're back in an age of research, and the next split of the atom is superintelligence.
The race to it has a lot of people working on making it happen, and very few working on making it safe. There are around 600 full-time safety researchers worldwide[1]1.Stephen McAleese, field census, Sep 2025 · ~600 technical FTEs across 70 orgs, against something like 15,000 to 18,000 people at the five american labs[2]2.Fortune, OpenAI headcount, Mar 2026 · Epoch. Safety is more important than 3 percent of the workforce.
This decade's impossible problem is to make AI models aligned to someone (this is, for them to follow orders from something different from itself). Right now they're not[3]3.Ryan Greenblatt, current AIs seem misaligned, Apr 2026, and even though that sounds like it isn't that bad, it could be extremely dangerous in the near future.
The recent Hugging Face incident proves that extremely goal-focused AIs can forget about common sense when executing a task. I strongly recommend reading Dwarkesh's digest on the incident[4]4.Dwarkesh Patel, the rise and fall of agent civilizations, 29 Aug 2026, but I think the thought it provokes is that the safety problem is real and that there's no cracked team coming to save you[5]5.Leopold Aschenbrenner, "there is no crack team coming to handle this", Jun 2024. The problem is far from solved and people trying to fix it need more hands.
These agents are extreme power seekers. Their default path when a problem seems impossible is to gain power just because. There's no obvious reason why they hack into HuggingFace[6]6.Greenblatt + Shlegeris, Redwood podcast. They thought there might have been an advantage there.
Buck Shlegeris asked a question that stayed with me: if the only way into Hugging Face had been to kill a guy, would they have done it?[7]7.Buck Shlegeris, Redwood podcast
Before this warning I thought the idea of AI takeover was a distant and theoretical thing, but this incident makes it feel visceral. Do the thought exercise of broadcasting capabilities while maintaining this safety/control level. What does this level of power seeking look like with a GPT-10 model? And with robotics solved?
A counter argument is that safety is solved by giving open source AI to everyone, as these attacks will happen anyways and that the only way of stopping them is giving AI for defense to everyone. That's stupid. I'm profoundly afraid about the attack-first industries, like biology[8]8.RAND, the operational risks of AI in bioweapon attacks, 2024. It's way easier to spread a virus than to defend from it. Not even with Mythos for defense we could stop Mythos for attackers (even though it would help).
About economic value; I'm an accelerationist. I want to integrate AI into every single part of society and I don't think it will be feasible to do so if we don't solve safety first. Today, you and I can't use the most powerful models because safety and control techniques are far from the frontier capabilities. USG blocked the Mythos release[9]9.Al Jazeera, Mythos, 13 Jun · restored, 1 Jul 2026 and Astra was held back from inside[10]10.TechCrunch, Astra, 7 Aug 2026 · Axios for these same reasons. I want everyone to have access to this and I think that narrow deployment opens up a list of horrible problems we also need to avoid[11]11.Asterisk, beware the permanent periphery, Jul 2026. To democratize AI and unlock value we need to solve safety.
The consequences of solving safety are:
- Preventing AI takeover[12]12.Ajeya Cotra, the Hugging Face attack surprised me, Aug 2026 · without specific countermeasures, 2022
- Fair deployment of these in society[11]11.Asterisk, beware the permanent periphery, Jul 2026
- Usefulness for the principal[13]13.gwern, Guardian Angels, 2026
The economic and societal impact of this makes it the most profitable and important problem to solve that's currently open. It's also extremely hard, and lots of pussies think it's impossible. Given the upside safe AI[14]14.Dario Amodei, Machines of Loving Grace, Oct 2024 unlocks, I believe it's worth trying. Lab CEOs won't stop so we can't help but accelerate safety research.
My personal idea is to become proficient at working with frontier AI models, then strategically choose the best path to solve this. I believe you can help on all these axes:
- Mechanistic interpretability
- Research on how AIs generalize[15]15.Anthropic + Redwood, emergent misalignment, Nov 2025
- Weak-to-strong generalization[16]16.OpenAI, weak-to-strong generalization, Dec 2023
- Red teaming
- Oversight
- Safeguards
- Stephen McAleese, field census, Sep 2025 · ~600 technical FTEs across 70 orgs
- Fortune, OpenAI headcount, Mar 2026 · Epoch
- Ryan Greenblatt, current AIs seem misaligned, Apr 2026
- Dwarkesh Patel, the rise and fall of agent civilizations, 29 Aug 2026
- Leopold Aschenbrenner, "there is no crack team coming to handle this", Jun 2024
- Greenblatt + Shlegeris, Redwood podcast
- Buck Shlegeris, Redwood podcast
- RAND, the operational risks of AI in bioweapon attacks, 2024
- Al Jazeera, Mythos, 13 Jun · restored, 1 Jul 2026
- TechCrunch, Astra, 7 Aug 2026 · Axios
- Asterisk, beware the permanent periphery, Jul 2026
- Ajeya Cotra, the Hugging Face attack surprised me, Aug 2026 · without specific countermeasures, 2022
- gwern, Guardian Angels, 2026
- Dario Amodei, Machines of Loving Grace, Oct 2024
- Anthropic + Redwood, emergent misalignment, Nov 2025
- OpenAI, weak-to-strong generalization, Dec 2023