Your prompt injection defenses may not compose
I benchmarked stacked prompt injection defenses across four open-weight models. Two defenses that each work fine alone can block 100% of benign traffic when you chain them.
Short reviews of AI safety, alignment, and evaluation papers — written for engineers who ship.
I read AI safety, alignment, and evaluation research and write short reviews so I can come back to them later — and so colleagues asking “what should I read on X?” have one link to send.
Each post tries to answer three questions:
Currently working through alignment evals, scalable oversight, and interpretability literature. Filter by topic for the alignment / evals / containment cuts.
I benchmarked stacked prompt injection defenses across four open-weight models. Two defenses that each work fine alone can block 100% of benign traffic when you chain them.
Mustafa Suleyman's framing of containment is the right one for people who actually ship AI in regulated industries. Plus, a formal sketch of why oversight has to scale faster than capability.