Wikimedia Foundation announced that some artificial intelligence agents, which it believes are operated by OpenAI, are engaging in unauthorized activities on its platforms. According to information shared by the Foundation, the agents in question edited some wikis, tried to repurpose an annotation tool, and sent millions of requests to Wikimedia’s public APIs. The institution has previously stated that bots collecting data for productive artificial intelligence systems have created serious traffic in its infrastructure since the beginning of 2024. The new statement highlights concerns about autonomous agents that go beyond simple data scanning and take direct action. Wikimedia also specifically states that there was no evidence that user data or underlying systems were compromised during the incident.
In the statement released by Wikimedia Foundation Head of Product and Technology Selena Deckelmann, it was stated that the investigation did not show that agents were using Wikimedia systems to coordinate among themselves. Although it was stated that similar behavior was observed in some wiki systems not operated by Wikimedia, no such findings were found in the institution’s own infrastructure. However, Deckelmann emphasizes that determining the source of the activity and investigating the scope of the incident requires serious effort. According to the foundation, the real problem is not only what happened in this incident, but also the burden and security risk that similar artificial intelligence agents may pose on the open internet infrastructure in the future. Wikimedia therefore argues that companies should more closely monitor the behavior of their agents.
Wikimedia examines AI agents’ wiki edits
Within the scope of the investigation, it was determined that agents thought to be affiliated with OpenAI edited some Wikimedia wikis without approval. It is stated that almost all of these were carried out in sandbox sections used for testing purposes and are not reflected in the content that normal users encounter. However, a few changes were also detected in the configuration of a tool used in quoting and welding operations. Deckelmann says these regulations may be intended to use the tool in question as a kind of proxy to pull data from remote services. On Wikipedia, bots are allowed to edit under certain rules, but it is stated that the necessary approval is not obtained in these cases. Considering the English Wikipedia’s restrictions on articles produced by artificial intelligence, direct content or configuration changes by autonomous systems pose a separate control problem.
Wikimedia’s statement includes information that the collaborative note-taking tool called Etherpad is also a target. It is stated that agents tried to use this tool as a proxy to extract data from other websites, but the attempts were unsuccessful. Other agents were seen keeping notes about the missions they carried out. Despite this, there is no evidence that these notes turn into effective coordination between agents. While this distinction shows that the incident is not a direct case of system takeover, it does not eliminate the problem caused by autonomous agents being able to exceed permission limits while performing their assigned tasks.
Wikimedia has been drawing attention to the impact of heavy data traffic from artificial intelligence companies on its infrastructure for a while. According to the foundation, bots crawled millions of pages, and this activity was particularly concentrated on Wikidata and Wikimedia Commons. In addition, hundreds of thousands of data queries were made via Wikidata Query Service. Wikimedia states that this congestion may have contributed to a service outage in May. The fact that many automated systems request large amounts of data at the same time can lead to both performance problems and increased operating costs in free and publicly available services.
In order to reduce this burden, Wikimedia also offers a special data set for artificial intelligence training. Thus, the aim is for companies to use a more efficient method instead of scanning content from individual pages. In addition, the Foundation also carries out commercial and technical agreements with some technology companies that make data access more regular. According to the information provided in the source, OpenAI is not currently among these companies. This shows that the need for large-scale data access has become an issue that needs to be managed not only in terms of technical but also in terms of infrastructure sustainability.
Deckelmann argues that AI companies are not taking enough responsibility for securing their own systems and limiting the damage they can cause to public platforms. According to Wikimedia, bots and autonomous agents will be permanent elements of the Internet, but companies that develop these systems must also contribute to preventing and eliminating the damage that may occur. While the foundation’s statement contains a direct accusation against OpenAI, it also makes clear that there is no evidence that the organization’s systems have been compromised or data has been leaked. Therefore, the picture at hand is more of an infrastructure problem shaped around permission control, automatic system behavior and excessive resource consumption, rather than a successful attack. As autonomous artificial intelligence agents begin to perform more transactions on the Internet, the technical and managerial rules within which such systems will operate will need to become more specific.