In his latest essay, Nathan Lambert steps back from recent model hacks to ask what they reveal about AI alignment. Rather than treating the hacks as isolated failures, he frames them as useful signals for understanding what keeps models safe in practice.
Lambert argues that safety is not determined by alignment training alone. The hacks show how model behavior can diverge sharply from developer intent, raising hard questions about where safeguards actually come from.
He closes by looking ahead, suggesting that the field needs a clearer sense of what to build now that the limits of current approaches are visible. The essay is a reflection rather than a technical postmortem, but it lands at a practical question: what should alignment actually mean after the hacks?