Guardrails For AI Didn’t Work. So Bring In The Philosophers
“Efforts [to stop the generation of harmful content] were initially focused around putting in simple black-and-white guardrails, such as forbidding a model from talking about bombs entirely. But these proved clumsy and easy to circumvent. Now, companies are p…
A friend sent me a text recently after reading about the latest OpenAI model that went rogue and launched a cyberattack on the tech platform Hugging Face in mid-July. Is this bad? he asked. Yes. W… [+4309 chars]