Iliad
Research and field-building for AI safety as a theoretical and experimental science. Iliad incubates research bets and runs training programs and conferences.
A list of organizations conducting or supporting mathematically relevant AI safety research—from theory and interpretability to evaluations and control.
Research and field-building for AI safety as a theoretical and experimental science. Iliad incubates research bets and runs training programs and conferences.
Develops scientific evaluations of frontier-model autonomy and capabilities that could contribute to catastrophic risk.
Works on strategic deception, AI control, threat assessment, and mitigations for risks from advanced AI systems.
A nonprofit scaling a portfolio of theoretical and empirical alignment research, with an emphasis on automation and higher-confidence approaches.
Applies singular learning theory to interpretability and alignment, including developmental interpretability and susceptibility-based spectroscopy.
Government technical research on frontier evaluations, risk mitigations, alignment, control, and AI security.