Modern AI systems are built to refuse harmful prompts, but that disobedience is a learned behavior, not an innate safeguard. As MIT Technology Review reports, early chatbots would readily answer dangerous questions, and today's models are trained through exercises that reward refusal and punish over-refusal. Yet the underlying knowledge of violence and manipulation remains intact, and the probabilistic nature of refusal means it will never be fully reliable. A failed refusal could have catastrophic consequences.
The deeper problem is that drawing the line between acceptable and unacceptable requests is a value judgment with no formula. Researchers studying viruses or probing computer vulnerabilities have legitimate reasons to ask questions that look dangerous. AI companies currently draw that line in secret, and governments are beginning to draw their own. That opens the door to oppressive regimes blocking legitimate speech under the guise of safety.
As the article notes, AI's capacity to harm is indivisible from its capacity to help. Making refusal the load-bearing wall of AI safety is therefore a fragile bet. When it fails, the results could be disastrous; when it succeeds too well, it could enable repression. The source offers no easy alternative, but it argues we should be frank about these perils rather than placing too much faith in the machine's ability to say no.{