News

Guardrails For AI Didn’t Work. So Bring In The Philosophers

  • The AI Humanist--Substack.com
  • published date: 2026-08-03 17:30:00 UTC

“Efforts [to stop the generation of harmful content] were initially focused around putting in simple black-and-white guardrails, such as forbidding a model from talking about bombs entirely. But these proved clumsy and easy to circumvent. Now, companies are p…

A friend sent me a text recently after reading about the latest OpenAI model that went rogue and launched a cyberattack on the tech platform Hugging Face in mid-July. Is this bad? he asked. Yes. W… [+4309 chars]