Institute for Responsible Superintelligence

RESI is a nonprofit research institute building the scientific foundations needed to make superintelligence safe by design. It is founded by researchers who have already turned seemingly murky concepts into precise definitions and usable mechanisms, in cryptography and other fields. The goal is to do the same for AI safety: safe-by-design means safety properties are specified in advance and achieved by mechanisms whose guarantees can be analyzed before deployment, rather than testing for safety problems post implementation.

Testing, benchmarking, and red-teaming are crucial for finding failures, but tests cannot establish that important classes of problems are absent. Across many fields, a deeper pattern has repeated. Notions that appeared to be inherently vague were given precise definitions, and unexpected inventions made those definitions usable for design. Safety shifted from something chased after the fact to a property that could be aimed for by construction.

Cryptography is a good example. Long ago, encryption schemes were ad hoc: constantly broken, red-teamed, and patched. Modern cryptography began when researchers formulated unexpected definitions of security and invented primitives (public-key encryption, digital signatures, and zero-knowledge proofs) that could satisfy them under explicit mathematical assumptions. Those inventions have withstood decades of attack and enormous increases in computing power. The same shift appears elsewhere. Early aviation relied on fly-fix-fly methods until phenomena such as turbulence could be modeled well enough to design against. Other enduring mechanisms, from auctions to contracts, can also achieve robustness by relying on incentives and laws we can understand.

RESI is a bet that this kind of shift is possible for superintelligence safety, and that it can happen by bringing together researchers who aim their efforts at responsible superintelligence, using frontier AI tools. RESI is founded by Turing Award winner and co-inventor of zero-knowledge proofs Shafi Goldwasser, renowned cryptographer Vinod Vaikuntanathan, and AI safety and ethics researcher Adam Tauman Kalai, who left OpenAI’s Safety Systems team to start the institute. Participating researchers will span computer science and other fields relevant to AI safety, including mathematics, economics and law, with an active visitor program. RESI is based in Cambridge, Massachusetts.

The concentration of talent is key. RESI aims to recreate the conditions which led to big ideas in the past, amplified by AI tools, with an open, mission-first culture that values sharing over being first to publish. Working groups will open new directions and revisit classical questions in light of superintelligence, while visitors will keep the institute connected to frontier labs and the scientific research community.

Mission

To build the foundations of safe-by-design superintelligence.

Outputs

RESI will develop an approach to AI safety that will remain relevant as intelligence increases. The approach calls for three types of outputs. First, we will map out which safety properties are meaningful, achievable under stated assumptions, or impossible. Second, we will develop mechanisms, protocols, and architectures that provide useful guarantees, and study whether those guarantees survive composition into larger systems involving models, tools, people, and institutions. As in modern cryptography, definitions and constructions develop together: a new protocol may reveal the right concept, and an impossibility result may show that a problem must be reformulated. Third, we will build implementations and proofs of concept of these constructions. Some constructions, such as harnesses, may be implemented at relatively low cost. Others may involve training entirely new frontier models, which requires resources well beyond what RESI expects to have initially. In that case, we will share proof-of-concept implementations so that frontier AI developers can evaluate and adopt them.

AI safety does not yet have a mature theory of definitions, constructions, composition, and impossibility comparable to modern cryptography. But early results already point in that direction: limits on inspection-based safety (including undetectable backdoors), definitions and mitigations for hallucinations as an incentive problem, undetectable watermarking of AI-generated content, and methods for boosting the safety of model ensembles under explicit assumptions. RESI exists to develop these seeds into a field aimed at superintelligence, where safety must be designed in rather than patched on afterward.

Harnessing the power of AI

Research at RESI is AI-assisted from the start. We will develop and continually improve harnesses that help theoretical alignment research using frontier agents throughout the research process, e.g., ideation, stress-testing definitions, writing and checking proofs, and turning theoretical ideas into experiments. These tools will let us explore new directions, iterate quickly, and tackle questions that might otherwise be out of reach.

Working Groups

Superintelligence safety has many aspects, and no single discipline sees the whole picture. RESI's working groups will open new directions and revisit many classical questions in light of superintelligence. Each working group will be organized by one or more leaders who will decide on its structure. The topics and working group leaders will be determined soon. Initial directions may include verification and delegation, modular architectures, incentives and strategic behavior, cryptographic mechanisms for AI safety, and legal mechanisms for superintelligence.

Culture

RESI’s culture is mission-first: open collaboration and early sharing of ideas, rather than waiting until results are posted. To protect junior researchers, we will actively support their credit and careers. Internal discussions at RESI follow the Chatham House Rule so participants can speak candidly while ideas can still be shared.

Founding Scientific Team

Shafi Goldwasser

Inventor of zero-knowledge proofs and other seminal strands of modern cryptography. Turing Award winner. Selected publications:

  • S. Ball, G. Gluch, S. Goldwasser, F. Kreuter, O. Reingold, G. Rothblum, Computational Barriers to Filtering for AI Alignment. ICLR 2026
  • N. Amit, S. Goldwasser, O. Paradise, G. Rothblum. Models That Prove Their Own Correctness. ICML 2024.
  • S. Goldwasser, M. Kim, V. Vaikuntanathan, and O. Zamir. Planting Undetectable Backdoors in Machine Learning Models. FOCS 2022.
Adam Tauman Kalai

AI safety and ethics researcher coming from 2.5 years at OpenAI. Works in learning theory, game theory, and fairness, and bridges frontier-AI practice with theory. Majulook Prize winner. Selected publications:

  • A. Kalai, O. Nachum, S. Vempala, and E. Zhang. Evaluating large language models for accuracy incentivizes hallucinations. Nature 2026.
  • A. Kalai, Y. Tauman Kalai, and O. Zamir. Consensus Sampling for Safer Generative AI. IASEAI 2026.
  • T. Eloundou, A. Beutel, D. Robinson, K. Gu-Lemberg, A. Brakman, P. Mishkin, M. Shah, J. Heidecke, L. Weng, and A. Kalai. First-Person Fairness in Chatbots. ICLR 2025.
Vinod Vaikuntanathan

MIT cryptographer working on fully homomorphic encryption, lattice-based cryptography, and secure computation. Gödel Prize winner. Selected publications:

  • V. Vaikuntanathan and O. Zamir. Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange. FOCS 2026.
  • S. Goldwasser, J. Shafer, N. Vafa, and V. Vaikuntanathan. Oblivious Defense in ML Models: Backdoor Removal without Detection. STOC 2025.
  • S. Goldwasser, M. Kim, V. Vaikuntanathan, and O. Zamir. Planting Undetectable Backdoors in Machine Learning Models. FOCS 2022.

Research Team

Ran Canetti

Boston University. Design, analysis, and composition of cryptographic protocols. Program obfuscation and applications. Selected publications:

  • R. Canetti, S. Goldwasser, O. Zamir. Proofs of Ownership for Machine Learning Models. arXiv 2026.
  • R. Canetti, C. Chamon, E. Mucciolo, A. Ruckenstein. Towards program obfuscation via local mixing. TCC 2024.
Miranda Christ

MIT postdoc, PhD from Columbia. Works on practically motivated cryptography, ML, and watermarking for AI-generated content. Selected publications:

  • M. Christ, N. Golowich, S. Gunn, A. Moitra, D. Wichs. Improved Pseudorandom Codes from Permuted Puzzles. STOC 2026.
  • M. Christ, S. Gunn. Pseudorandom Error-Correcting Codes. CRYPTO 2024.
Yannai A. Gonczarowski

Harvard economist and computer scientist working at the intersection of economic theory and theoretical computer science, including mechanism design and market design. ACM SIGecom Doctoral Dissertation Award winner. Selected publications:

  • S. Fish, Y.A. Gonczarowski, R.I. Shorrer. Algorithmic Collusion by Large Language Models. EC 2026.
  • Y.A. Gonczarowski S.M. Weinberg. The Sample Complexity of Up-to-ε Multi-dimensional Revenue Maximization. FOCS 2018 / JACM 2021.
Sam Gunn

PhD from UC Berkeley. Works on cryptography and foundations of AI. Selected publications:

  • S. Gunn. How to sketch a learning algorithm. 2026.
  • S. Gunn, X. Zhao, D. Song. An Undetectable Watermark for Generative Image Models. ICLR 2025.
Yael Tauman Kalai

MIT cryptographer and theoretical computer scientist. Works on verifiable delegation, interactive proofs, and cryptographic protocols. Winner of the ACM Prize in Computing. Selected publications:

  • L. Chen, Y. Tauman Kalai, Z. Xi. How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs. ICML 2026.
  • S. Goldwasser, Y. Tauman Kalai, G. Rothblum. Delegating Computation: Interactive Proofs for Muggles. STOC 2008.
Noam Kolt

Assistant Professor at the Hebrew University Faculty of Law and School of Computer Science and Engineering. Selected publications:

  • N. Kolt, Superintelligence and Law. Harvard Journal of Law & Technology 2026.
  • N. Kolt et al., Legal Alignment for Safe and Ethical AI. TMLR 2026.
Daniela Rus

MIT roboticist and director of CSAIL working on autonomous systems, intelligence, and safe control and planning for embodied AI. MacArthur Fellow. Selected publications:

  • W. Xiao, D. Rus, et al. BarrierNet: Differentiable Control Barrier Functions for Learning of Safe Robot Control. IEEE Transactions on Robotics 2023.
  • W. Xiao, J. Wang, C. Gan, R. Hasani, M. Lechner, and D. Rus. SafeDiffuser: Safe Planning with Diffusion Probabilistic Models. ICLR 2025.
Jonathan Shafer

MIT postdoc and incoming professor at the Weizmann Institute of Science, PhD from UC Berkeley. Works on learning theory and its connections to cryptography. Selected publications:

  • Z. Chase, S. Hanneke, S. Moran, and J. Shafer. Optimal Mistake Bounds for Transductive Online Learning. NeurIPS 2025 (Best Paper Runner-Up).
  • S. Goldwasser, G. Rothblum, J. Shafer, and A. Yehudayoff. Interactive Proofs for Verifying Machine Learning. ITCS 2021.
Yaron Singer

CEO and co-founder of Frontier Security. Previously VP at Cisco, CEO and founder of AI Security startup Robust Intelligence (acquired by Cisco), and Harvard professor of Computer Science and Applied Mathematics. Works on AI in cybersecurity, AI security, adversarial robustness, and robust optimization. Sloan Fellow. Selected publications:

  • P. Kassianik, B. Nelson, and Y. Singer. Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents. CAMLIS 2026.
  • A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. S. Anderson, Y. Singer, and A. Karbasi. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. NeurIPS 2024.
Andrew V. Sutherland

MIT mathematician working in AI-assisted mathematical discovery, formal verification, large-scale mathematical databases, computational number theory, and arithmetic geometry. Selected publications:

  • J. S. Ellenberg, C. S. Fraser-Taliente, T. R. Harvey, K. Srivastava, and A. V. Sutherland. Generative Modelling for Mathematical Discovery. Advances in Theoretical and Mathematical Physics 2025.
  • W. Sawin and A. V. Sutherland. Murmurations for Elliptic Curves Ordered by Height. 2025.
Mirac Suzgun

PhD candidate in CS at Stanford and JD from Stanford Law School. Works on LLM reasoning, safety, and factuality, with particular interests in understanding and mitigating hallucinations and evaluating AI systems deployed in high-stakes domains such as law and medicine. Selected publications:

  • M. Suzgun, M. Yuksekgonul, F. Bianchi, D. Jurafsky, J. Zou. Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory. EACL 2026.
  • M. Suzgun, T. Gur, F. Bianchi, D.E. Ho, T. Icard, D. Jurafsky, J. Zou. Language Models Cannot Reliably Distinguish Belief from Knowledge and Fact. Nature Machine Intelligence 2025.
Neekon J. Vafa

Junior Fellow of the Harvard Society of Fellows, PhD from MIT. Works on cryptography and its connections to statistics and trustworthy AI. Selected publications:

  • A. Bogdanov, A. Rosen, N. Vafa. Statistically Undetectable Backdoors in Deep Neural Networks. ICML 2026.
  • A. Bogdanov, A. Rosen, N. Vafa, and V. Vaikuntanathan. Adaptive Robustness of Hypergrid Johnson-Lindenstrauss. STOC 2026.
Rebecca Wexler

Alfred W. Bressler Professor of Law at Columbia. Works on evidence and criminal law, including interactions of technological and legal design for secrecy, authentication, and accountability. Selected publications:

  • S. Barrington, E. Cooper, H. Farid, R. Wexler. AI-Generated Voice Evidence Poses Dangers in Court. Lawfare 2025.
  • D. Bitan, R. Canetti, S. Goldwasser, R. Wexler. Using Zero-Knowledge to Reconcile Law Enforcement Secrecy and Fair Trial. ACM CSLaw 2022.

Operations

Emily Uyeda Kantrim

Emily Uyeda Kantrim’s work spans nonprofit, government, research, and social innovation, with a focus on turning ambitious ideas into durable institutions, scalable systems, and measurable public impact. From ideation to implementation, Emily has turned hand-sketched product designs into open-access tools for creation and entrepreneurship across six continents, scaled mobile public health operations from $4M to $93M in two years, and recognized an opening in municipal policy and built it into an operational social-service model that reached national adoption.

Planned visitors

Affiliated groups

FAQ

Funding

RESI is fiscally sponsored by the Edward Charles Foundation, a 501(c)(3) public charity (3% overhead). RESI is funded through philanthropic donations, and our list of funding supporters will be announced shortly. Individuals can give at every.org/resi. Contributions of cash, stock, and other assets are tax-deductible to the extent permitted by law. Donations will enable us to bring in more researchers (especially postdocs and junior researchers) and further accelerate our research using AI. To donate or discuss other forms of support, including compute credits or infrastructure, contact funding@resi.org.