Overview

Alignment is not a philosophy seminar here. It is measurement. You will build a benchmark that can actually fail, write a rubric a model can apply and then check it against independent graders, push a model toward and away from a set of values by prompt and by fine-tune, and prove the change was real rather than a story you told yourself afterwards. The discussion threads — reward misspecification, scalable oversight, hallucination and sleeper agents, and whose values we are aligning to in the first place — run alongside the work rather than instead of it.

Where This Class Leads

What you leave able to do

  • design an interaction protocol for responsible use
  • arbitrate between the measure a system optimises and the goal
  • arbitrate which internal feature drives a models behaviour
  • govern a swarm by information hierarchy
  • select among operating boundaries for inputs a system was not built for
  • synthesise an accountability regime for an unattended agent
  • originate oversight that holds when the system outpaces the reviewer

How you show it. The scheme, a run where output exceeded what a reviewer could read, and an error the scheme surfaced anyway.

Join Us

Want to schedule this class for your team?

Contact us: liz@themultiverse.school