AI Safety & Paper Reviews

Short reviews of AI safety, alignment, and evaluation papers — written for engineers who ship.

I read AI safety, alignment, and evaluation research and write short reviews so I can come back to them later — and so colleagues asking “what should I read on X?” have one link to send.

Each post tries to answer three questions:

  1. What is the paper actually claiming?
  2. Why should an engineer (not a researcher) care?
  3. Where would I push back?

Currently working through alignment evals, scalable oversight, and interpretability literature. Filter by topic for the alignment / evals / containment cuts.

2026

The Coming Wave — a working engineer's review

Mustafa Suleyman's framing of containment is the right one for people who actually ship AI in regulated industries. Plus, a formal sketch of why oversight has to scale faster than capability.