User-Aligned Is Not User-Safe

A model that perfectly follows one user's intent can still be dangerous to everyone around that user.

  • #ai
  • #governance
  • #security
Personal intent constrained by responsibility rings

“Aligned with the user” sounds great until you ask which user.

That is the useful tension in TechCrunch’s write-up of George Hotz’s argument for locally controlled, user-aligned AI. The appealing version is obvious: a personal model that works for you, runs closer to you, and does not bend every request through a giant platform’s policy stack. There is a real product dream in that.

Then the thought experiment gets spicy: what if the user wants help doing something destructive?

This is where “alignment” stops being a magic word. An AI system can be aligned with the person typing and misaligned with the spouse, coworker, child, customer, bystander, regulator, or future investigator who is also inside the blast radius. Product safety is not just about the current session. It is about the network of people a capability can touch.

That same shift shows up in a much less dramatic, more practical way as OpenAI starts hiring for family, caregiver, and older-adult experiences. A household product is not one user with one preference file. It is parents, teens, caregivers, older adults, shared devices, private questions, crisis moments, and people who need different defaults.

For builders, the lesson is not “local bad, cloud good” or the reverse. It is that user intent is only one input to the product contract.

The better design question is: who can be harmed if this request succeeds?

If you cannot answer that, your alignment story is probably just customer success wearing a lab coat.

All notes · RSS