Our mission is to build aligned superintelligence

A safe transition to a post-AGI world requires the technical ability to steer AGI, effort to steer it towards the right objectives, and society-wide robustness against adversarial actors using AGI. Quite uncomfortably, nobody knows how to reliably achieve these outcomes today.

Here, we summarize our thinking on three priority subproblems we want to help address:

  1. AI alignment

    Alignment is hard

    Solving the alignment problem is difficult. Even scoring 100% on an alignment eval doesn’t guarantee that a model is aligned (it could be faking it) or that the techniques will work on stronger models. Monitoring isn’t a substitute for alignment training either. And even perfect interpretability of frozen weights wouldn’t be sufficient to ensure the model doesn’t cause mayhem during RL training. It’s unclear what a theoretically sound solution to alignment looks like and most current lines of research seem unlikely to converge to one.

  2. Correlation between the capitalist race and safety

    Marrying the capitalist race and safety

    Capitalism is a wonderful optimizer for efficiency and progress, and it could be a force for solving technical safety problems if we build a race track that requires companies to do so. Many companies adopted voluntary safety policies, but such frameworks would be most impactful if written into law, internationally. Coordination on alignment research and gating AI progress by safety standards would benefit everyone and it seems plausibly achievable that the US and China could agree on terms.

  3. AI-assisted alignment R&D

    AI-assisted Alignment R&D

    If robustly aligning superintelligence with human intent were easy, it’d be solved by now. We may want to use AI to autonomously design, develop and evaluate new alignment techniques. To do so, we need “sufficient alignment techniques” to align that system. On the initial iteration, a shortlist of approaches can be reviewed by human researchers, and access to the internet could be disallowed, allowing us to make use of imperfectly or narrowly aligned superhuman systems. From there, the aligned model could iteratively develop and align stronger models autonomously.


If you have thoughts on our approach, we would be delighted to hear from you at safety@magic.dev and if our priorities resonate with you, we invite you to apply to join us.