Resources

What we send to someone who asks where to start. Almost all of it is in English, which is where the field publishes.

Five readings to start with

In this order. The first two assume nothing; the last one is a sequence and takes several sittings.

Videos

To get a feel for the field without reading a whole article.

  • The reference video library on alignment, control and the problems that are still open.

  • AI In Context

    80,000 Hours

    Long documentaries on where AI is heading. The first one, on the AI 2027 scenario, passed ten million views.

  • Animations adapting classic essays on AI, rationality and existential risk.

  • Short animated explainers, each on one concrete risk or proposal.

  • Doom Debates

    Liron Shapira

    Debates and interviews where both positions on the risk are argued face to face.

Podcasts

Long conversations, good for a commute or the gym.

  • AXRP

    Daniel Filan

    The reference podcast on technical alignment: long interviews with the people doing the research.

  • Conversations with researchers, founders and policy people about what to do with your own career.

  • Weekly, with people from the frontier labs. Useful for keeping up with what comes out.

Books

For when you want the whole argument and not a summary.

  • The Alignment Problem

    Brian Christian

    An accessible account of why it is hard to align a system with what we actually want.

  • Human Compatible

    Stuart Russell

    The case, from one of the authors of the classic AI textbook, for rebuilding the field around uncertainty.

  • Superintelligence

    Nick Bostrom

    The book that brought the discussion of general AI into public debate. It is from 2014 and it shows, but it fixed the vocabulary.

Courses

With deadlines, assignments and someone on the other side.

  • AGI Strategy

    BlueDot Impact

    The entry point we recommend most. You work in groups of eight with a facilitator, taking a threat apart step by step to see which of those steps is worth intervening on.

  • Every unit leaves something finished: a brief aimed at someone who decides, a map of who has authority over frontier models, and a position of your own defended in writing.

  • Technical AI Safety

    BlueDot Impact

    It runs through alignment, interpretability, evaluations and control so you can tell which one fits you. Whoever finishes can apply to BlueDot's project sprint.

  • AGI Safety Course

    Google DeepMind

    Seventy-five minutes of recorded talks with exercises. It separates two ways an objective goes wrong: the model gaming the criterion it is rewarded on, or the model learning a different objective.

  • AI Safety, Ethics and Society

    Center for AI Safety

    It treats safety as an engineering problem and not only a machine learning one: it carries over to AI what aviation and the nuclear industry learned about accidents.

  • ARENA

    Five weeks, for people who already code

    You write code from day one: train a transformer, open it up to see what it computes inside, and build a test that measures a model.

If you want to see the whole field

aisafety.com is an open directory with hundreds of organisations, programmes and projects, and with who funds each one.

They go further with company

Go through them with someone else

The reading group discusses one every fortnight, and the WhatsApp group is where you ask what does not add up.