All thesis threads

Cumulative thesis

Safety and Alignment

How do we keep capable AI accountable to human judgment?

Alignment asks whose values AI serves. Safety asks what happens when it fails. This thread connects model interpretability, the risks of autonomous action, and the human judgment we need to preserve.

Where we land

More capable AI should remain open to inspection, challenge, and correction. Aligning the machine means preserving the human capacity to question it.

Essays in this thread

3 essays
  1. Alignment and Critical Thinking essay image

    Alignment and Critical Thinking

    AI alignment is only half the problem. The critical thinking we are offloading is the very faculty alignment depends on.

    Essays
  2. The Internet Is Not Ready essay image

    The Internet Is Not Ready

    General intelligence is lowering the cost of attack and putting machine-speed defense within reach.

    Essays
  3. What the Model Thinks But Doesn't Say essay image

    What the Model Thinks But Doesn't Say

    Anthropic found a global workspace inside Claude: a silent, evolving set of unspoken words the model thinks with. Learning to read it may matter more for the agent economy than the next benchmark record.

    Essays