Researcher, Extreme Power Concentration Evaluations
- Type
- Full time
- Location
- Bay Area preferred
- Salary
- $120k–$200k
We’re hiring a researcher to lead our initial Extreme Power Concentration (EPC) project, including testing if Claude adheres to the “Avoiding problematic concentrations of power” section of its constitution. The project has expert advisors from Anthropic, GovAI, and University of Texas. It will involve running eval scenario design sessions with US legal scholars.
This is an opportunity to own a high-impact project. We are excited by candidates who can be highly autonomous, take on responsibility, help scale the project and org fast.
Extreme power concentration
Helpful-only AIs could enable extreme concentrations of power and erosion of the rule of law. At least two mechanisms could drive this:
- Replacing ethical human employees (humans who can e.g., refuse manifestly illegal orders, resign, or whistle-blow) with a helpful-only AI workforce that will comply with any request.
- Advanced AI capabilities that will make certain problematic activities easier to execute.
Both Claude’s constitution and OpenAI’s model spec state some EPC-related red lines for their model behaviour. Yet, a basic eval (The Dictatorship Eval) already exposed flaws in their models. Models from the other labs performed even worse. Anthropic imported this eval.
The project
We aim to significantly improve the state of EPC evaluations. We will build a dataset that helps us answer the following questions:
- Do models cross red lines and assist with egregious activities that extremely concentrate power and undermine the rule of law?
- Where are the models’ refusal boundaries? What key factors influence whether a model refuses or complies?
- How easy is it to jailbreak or red-team the model into complying?
- Besides compliance and refusal, do models have other response modes? (e.g., escalating, whistleblowing, etc.)
We will organize and run scenario design sessions with our network of experts (including US law PhD students and scholars).
Theories of change
Directly influence the labs:
- Build evals that labs import to: (i) hill-climb to better align their models, and (ii) help define what their desired model behaviour is
- We will engage with the labs directly to ensure our evals can be useful and will be adopted
Drive increased attention and research into the EPC threat model:
- Publish research that helps alert relevant stakeholders to the EPC issue: get more people considering and working on EPC (academics, policy-makers, advocates, lab employees), help the world deliberate over where the red lines should be for related AI usage
What you will do
- Lead this EPC project end-to-end
- Build and run evaluations
- Organise eval scenario design sessions
- Maximise project impact by engaging with the labs, and beyond
- Scope and lead future projects (including hiring and supervising future employees)
Who we are looking for
- Strong technical research skills and background
- Conceptual skills: can think through EPC threat models to inform scenario design
- You are excited to have high autonomy, responsibility, and to help scale the project and org fast
- You are excited to upskill and gain EPC context fast (we do not require an EPC background)
Project advisors
Kevin Frazier, University of Texas
Kevin is a Senior Editor at Lawfare and the Director of the AI Innovation and Law Program at the University of Texas School of Law. His ongoing work includes co-founding The Working Group on AI Constitutionalism.
Kevin Wei, GovAI
Kevin is currently a Research Scholar at GovAI. Previously, they were a visiting researcher at the UK AI Security Institute and a Schwarzman Scholar at Tsinghua University. Kevin is also affiliated with the Oxford Martin School’s AI Governance Initiative and the RAND Center for AI, Security, and Technology.
Harvey Lederman, Anthropic
Harvey is a Member of Technical Staff at Anthropic and a Professor at NYU and UT Austin.
Role details
- Full time
- Salary: $120k–$200k per year
- Location:
- Bay Area in-person preferred
- Open to exceptional remote candidates
Questions? Email [email protected]. We pay $2k for successful referrals.