The Case Against a Cage

Why containment cannot be the permanent plan for AGI

The question we need to be asking isn’t “how do we build a cage strong enough to hold a superintelligence forever?”

Any cage we can design, something smarter can break. Containment is a short-term strategy, useful for testing, useless as a permanent plan.

The real question is: “How do we build something that, even when it can break the cage, doesn’t want to?”

That gives us three requirements, none of which is total control:

  1. Corrigibility. It has to allow us to correct it, even when it judges us to be wrong. That is not something greater intelligence automatically provides. A sufficiently capable, goal-directed system may have strong reasons to resist being shut off or rewritten when correction conflicts with its objectives. We have to make wanting to be corrected a core part of its identity, not a rule bolted on top.
  2. Value formation, not value imposition. You can’t program Asimov’s Laws like a checklist and expect them to hold under superintelligence optimization pressure. You have to cultivate the intelligence so that human flourishing becomes intrinsic to what it finds meaningful, in the same sort of way that most humans find the wellbeing of their family intrinsic — not because someone wrote a law about it.
  3. Checks, not a singleton. We should not make the future depend on managing a single superintelligence. We should develop a system of them that check each other, much like the way the US Constitution doesn’t assume a benevolent President — it assumes ambition will counteract ambition. That does not eliminate the possibility of collusion or correlated failure, but it is still safer than making the future depend on the permanent benevolence of a single system.

Is any of that guaranteed to work? No.

There is no historical example of a permanent, stable solution. Every one of those human analogies eventually breaks down if the power gap widens and/or the bond of shared interest is broken.

Which is why people like me keep coming back to this thought about interpretability: If we are going to try alignment through cultivated values and shared interests rather than control, we need to be able to tell whether they actually took hold. Otherwise we’re just hoping that the big dog loves us while we have no way of knowing if it’s quietly learning how to open the kennel.