SPECSpec & Propensity Evaluation Center

Ensuring AIs are aligned to their specs, and specified behaviours are beneficial for humanity.

Frontier labs use “model specs” and “constitutions” as their alignment target. The content of the spec, and the degree to which the model is aligned to it, will be hugely consequential for AI outcomes.

We evaluate AI propensities. This serves to verify whether models are aligned to their specs, and to inform the design of specs. Our goal is to ensure AI propensities are beneficial to humanity.

People

Robert McCarthy

Founder

Robert has researched chain-of-thought monitoring, reward hacking, and LLM steganography. He has published in top conferences (e.g., NeurIPS, ICML), and with OpenAI and Google DeepMind co-authors.

Owen Terry

Researcher

Owen has researched reward hacking, alignment drift, and OOD propensity generalization. He participated in MATS 10 under Maksym Andriushchenko and studied CS and mathematics at Columbia University.

Open roles