Skip to content
Make AI Good

Graph · Person

Jan Leike

01 · In focus

One person, in the field.

The structured facts the source records about Jan Leike, the count of declared adjacencies in the corpus, and the federation map zoomed on this node and its neighbours.

person

0 declared connections

Kind
Person
Status
active
Confidence
high
Entity ID
person-jan-leike
Network
View in network

Tags germany, ai-safety, ai-alignment, rlhf, frontier-ai, openai, anthropic, deepmind, superalignment, whistleblower, researcher

Jan Leike · 0 direct neighbours visible

03 · Background

From the source record.

Body prose as it appears in movement-graph’s published markdown for this entity. Links to other corpus entities resolve to their graph page; links to deeper repo paths are kept as text so the page does not invent a route.

German AI alignment researcher; currently leads the Alignment Science team at Anthropic, working on scalable oversight, weak-to-strong generalization, and automated alignment research. He holds a PhD in machine learning from the Australian National University (supervised by Marcus Hutter) and completed a postdoctoral fellowship at the Future of Humanity Institute.

From 2016 to 2021 he worked in DeepMind's AI safety group, where he co-authored the 2017 paper "Deep Reinforcement Learning from Human Preferences" — early empirical groundwork for reinforcement learning from human feedback (RLHF). He joined OpenAI in 2021, leading the alignment team and contributing to InstructGPT (2022), the instruction-following model that brought RLHF into large-scale language models.

In June 2023 OpenAI named Leike and Ilya Sutskever co-heads of a newly formed Superalignment initiative, announcing a commitment of 20% of compute and a four-year timeline to solve alignment for superhuman AI systems. On 17 May 2024 he resigned — two days after Sutskever's own departure — publishing a thread on X stating that "safety culture and processes have taken a backseat to shiny products." OpenAI subsequently confirmed the Superalignment team's dissolution. The departure is documented in the insider whistleblowing from frontier labs strategy entry as one of the precipitating events behind OpenAI's pre-emptive reversal of its nondisparagement agreements and the Right to Warn open letter of 4 June 2024. He joined Anthropic later that month. He was named to TIME's list of the 100 most influential people in AI in both 2023 and 2024.

04 · Sources

Where this came from.

5 sources listed from the pinned corpus. Links are shown only when the source URL is a valid HTTP(S) address.

  1. x.com

    Checked 2026-06-12

    Leike's own X thread (17 May 2024) — primary source for his resignation from OpenAI and his stated reason that safety culture had taken a backseat to shiny products

  2. fortune.com

    Checked 2026-06-12

    Fortune (17 May 2024) — primary source for his role as Superalignment team co-head and public resignation statement

  3. cnbc.com

    Checked 2026-06-12

    CNBC (28 May 2024) — primary source for his move to Anthropic and stated research focus on scalable oversight and weak-to-strong generalization

  4. en.wikipedia.org

    Checked 2026-06-12

    Wikipedia article — corroborates DeepMind career (2016–2021), OpenAI role as Superalignment co-head, PhD at Australian National University under Marcus Hutter, and TIME 100 AI listings (2023 and 2024)

  5. jan.leike.name

    Checked 2026-06-12

    Personal website — primary source for current role leading the Alignment Science team at Anthropic and research focus areas

Source: entities/persons/person-jan-leike.md — movement-graph pin 5d136ad.