Google and researchers at Google DeepMind announced a new formation called DeepMind Institute, which aims to bring discussions about the future of artificial general intelligence to a broader platform. The institute’s management team includes DeepMind co-founder Shane Legg, Google executive James Manyika and Google DeepMind President Demis Hassabis. Legg will also serve as the organization’s executive editor. DeepMind Institute aims to open space for different approaches emerging in the global research community on artificial general intelligence, rather than just sharing the views shaped within Google and Google DeepMind. The founding announcement specifically states that contributing researchers may not agree on everything and that current views may change as new data comes from rapidly advancing artificial intelligence research.
The institute’s first publication consists of four separate articles, each of which focuses on different problems that artificial general intelligence may pose. Topics covered include policies that can be implemented against AGI-induced economic transformations, preserving the human-readable reasoning processes of AI models, principles to support human development and well-being, and a framework for how the most advanced AI models can be evaluated. This choice of topic shows that the debate is not reduced solely to the technical capacity of the models. Ranging from economic impacts to safety oversight, the broad framework addresses issues that companies, researchers, and regulators may face as AGI evolves. In addition, instead of presenting definitive policy prescriptions in published texts, it is preferred to discuss the advantages and limitations of different options.
DeepMind Institute brings the auditability of artificial intelligence models to the agenda
One of the articles written by Google DeepMind security researchers Rohin Shah and Anca Dragan focuses on the problem that it may become increasingly difficult to understand how advanced artificial intelligence models think and what intermediate steps they follow. The researchers argue that the shrinking window of transparency in models’ reasoning processes is not an inevitable outcome. While new architectures create more powerful models, it can also become more difficult to observe and control the internal processes of these systems. According to Shah and Dragan, developers and regulators need to clearly consider the tradeoffs between security and model capability at this point. The authors discuss options such as limiting the level of “opaque serial depth,” which is the amount of sequential calculations performed without producing a human-readable trail of reasoning, or requiring companies that develop less transparent models to prove that these systems are equally auditable.
In another article written by Demis Hassabis, a US-based standards body is proposed for the independent evaluation of the most advanced artificial intelligence models. According to Hassabis’ framework, companies can voluntarily submit high-end models they develop in the first stage for evaluation up to 30 days before launching them. Once the evaluation system is shown to be reliable and effective, passing these tests could be made one of the conditions for introducing advanced artificial intelligence models in the United States. Such a system could enable companies to establish common benchmarks beyond their own security tests. However, details such as the scope of independent evaluations, which models will be subject to this process, and how trade secrets will be protected will be decisive in the feasibility of such an approach.
It is envisaged that the proposed institution will design its tests together with artificial intelligence companies in the first stage, and then prepare its own independent and previously undisclosed evaluations in the later period. The main purpose of these tests, called “held-out” in the article, is to prevent developers from specifically optimizing their models according to previously known criteria. In this way, the real capacity and security vulnerabilities of a model can be examined under conditions where companies do not know the test content in advance. Hassabis also states that if the risks increase, the evaluation framework can be made more stringent. At the most advanced stage of this approach, if security measures lag behind technological development, it is also an option for leading artificial intelligence developers to slow down their work in a coordinated manner.
The DeepMind Institute’s first papers are being published at a time when the AI security debate is moving from general risk warnings to more concrete auditing and evaluation recommendations. Recently, options such as model developers disclosing more information, allowing external organizations to examine the systems, and slowing down the development speed in cases where security mechanisms are not sufficient have been brought to the agenda more frequently in the industry. According to the source text, this debate has become more visible in recent days, with Anthropic CEO Dario Amodei’s call for a controlled adjustment of the pace of development of the most advanced artificial intelligence systems, gaining support from some names in the industry. Despite this, it is not yet clear how to strike a balance between companies’ voluntary security commitments and binding regulations. As of its first publications, DeepMind Institute aims to create a discussion ground where different approaches ranging from technical security to economic policies can be compared, rather than presenting a single institutional view on AGI.