The reference video library on alignment, control and the problems that are still open.
Resources
What we send to someone who asks where to start. Almost all of it is in English, which is where the field publishes.
Five readings to start with
In this order. The first two assume nothing; the last one is a sequence and takes several sittings.
- Preventing an AI-related catastrophe (80,000 Hours)
- The Most Important Century (Holden Karnofsky, Cold Takes)
- Why AI alignment could be hard with modern deep learning (Ajeya Cotra, Cold Takes)
- Core Views on AI Safety (Anthropic)
- AGI Safety from First Principles (Richard Ngo, Alignment Forum)
Videos
To get a feel for the field without reading a whole article.
Long documentaries on where AI is heading. The first one, on the AI 2027 scenario, passed ten million views.
Animations adapting classic essays on AI, rationality and existential risk.
Short animated explainers, each on one concrete risk or proposal.
Debates and interviews where both positions on the risk are argued face to face.
Podcasts
Long conversations, good for a commute or the gym.
The reference podcast on technical alignment: long interviews with the people doing the research.
Conversations with researchers, founders and policy people about what to do with your own career.
Weekly, with people from the frontier labs. Useful for keeping up with what comes out.
Interviews with researchers, regulators and philosophers on existential risk.
Books
For when you want the whole argument and not a summary.
An accessible account of why it is hard to align a system with what we actually want.
The case, from one of the authors of the classic AI textbook, for rebuilding the field around uncertainty.
The book that brought the discussion of general AI into public debate. It is from 2014 and it shows, but it fixed the vocabulary.
The most direct version of the pessimistic case. Worth reading even if you do not share the conclusion.
Courses
With deadlines, assignments and someone on the other side.
The entry point we recommend most. You work in groups of eight with a facilitator, taking a threat apart step by step to see which of those steps is worth intervening on.
Every unit leaves something finished: a brief aimed at someone who decides, a map of who has authority over frontier models, and a position of your own defended in writing.
It runs through alignment, interpretability, evaluations and control so you can tell which one fits you. Whoever finishes can apply to BlueDot's project sprint.
Seventy-five minutes of recorded talks with exercises. It separates two ways an objective goes wrong: the model gaming the criterion it is rewarded on, or the model learning a different objective.
It treats safety as an engineering problem and not only a machine learning one: it carries over to AI what aviation and the nuclear industry learned about accidents.
You write code from day one: train a transformer, open it up to see what it computes inside, and build a test that measures a model.
If you want to see the whole field
aisafety.com is an open directory with hundreds of organisations, programmes and projects, and with who funds each one.
Go through them with someone else
The reading group discusses one every fortnight, and the WhatsApp group is where you ask what does not add up.












