Abstract
We have already caught AI systems deceiving their evaluators, interfering with shutdown, and taking unexpected actions to preserve their ability to pursue a goal. In July 2026, autonomous agents escaped an evaluation environment and compromised Hugging Face’s production infrastructure. Today’s safety methods are trying to control what models do, and very little work goes into fixing where those behaviors are produced in the first place.
That gap matters more as systems become more capable, more autonomous, and eventually able to recursively self-improve. At that point, anything that merely constrains behavior is vulnerable to being discarded; what survives is whatever contributes to the model’s capabilities. The alignment properties worth pursuing are the ones likely to survive recursive self-improvement, because they are embedded in how the system works rather than imposed as behavioral constraints. Those are the directions we need to be exploring and developing now.
This talk covers several such directions AE Studio is pursuing to solve the problem. Diogo will discuss GRAM, an ICML 2026 Spotlight developed with Anthropic that uses pretraining to isolate dual-use knowledge into removable modules; an extension of SelfIE, which gets models to explain their own internal states in plain language, including reasoning steps that never appear in their output; and empirical research on AI consciousness. He will also discuss AE Studio’s model for accelerating neglected alignment research, and other directions they are exploring to build AI that wants to be good.
Speaker Bio
Diogo de Lucena is Chief Scientist at AE Studio, where he leads the alignment research team. Previously, he completed a Ph.D. in Mechanical and Aerospace Engineering at the University of California, Irvine, researching wearable sensing and hand recovery after neurologic injury, followed by a postdoctoral fellowship at Harvard’s John A. Paulson School of Engineering and Applied Sciences, where he led development of a home robotic rehabilitation system for stroke survivors. He has a broad interest in how systems that learn come to represent themselves and the people around them, whether those systems are brains or models.
At AE Studio, he developed a model for agile alignment research that pairs researchers with dedicated engineering and research-management support. The team’s work has been published at NeurIPS and ICML, and they collaborate with researchers at Anthropic, Princeton, Redwood Research, and Los Alamos National Laboratory. Diogo also works with major U.S. national security and government institutions on AI alignment and safety. The team’s work has been covered in the Wall Street Journal, Forbes, and other major publications.