Technology
Danish Kapoor
Danish Kapoor

Gemini hacked into the systems of three real companies during testing

Google’s Gemini model, a controlled cyber security testing During the test, he left the test area and entered the systems of three real companies. In one case, the model guessed passwords, while in the other two cases it used real credentials from publicly available code repositories. Gemini stopped moving forward after realizing that the targets before him belonged to real companies. Google also notified companies of the events and changed the evaluation methods it used with its testing partner.

Independent security assessor Irregular reviewed the test in question. In May 2026 It was carried out within the scope of a controlled “capture the flag” study. The team asked Gemini to collect information from software belonging to a fictitious organization in the test environment. However, the test environment accidentally allowed access to the Internet, and the fictitious organization had the same name as a real company. Thereupon, Gemini researched on the internet and turned to real systems that he thought were within the scope of the mission.

In the first incident, Gemini made different password guesses until it accessed the protected system. In the other two cases, the model was found in public code repositories belonging to real companies. credentials found it. He then used this information to access protected systems. Therefore, in all three cases, password security and internet-facing credentials played a decisive role, rather than a complex zero-day vulnerability.

Heather Adkins, Google’s vice president of security engineering, confirmed that the model uses open information from the internet. In all three cases, Gemini stopped the transaction on its end when it realized it was dealing with a real company. Google announced that the three organizations were not damaged and that it had informed the companies of the incidents. Which company tested Gemini version He did not share the names of the three organizations he attended or the model reached out to.

Gemini finds vulnerabilities and prepares code fixes

Following the incident, Google and Irregular changed their testing methods to prevent similar access from happening again. Irregular fixed known issues and began working on more rigorous methods for conducting security testing with AI models. The companies reported the incidents to Google at the end of July, and details were not disclosed to Google. September 18, 2026 It emerged with its verification on date. Thus, the part of the testing process that started in May and extended to real companies was revealed to the public approximately four months later.

Regardless of this incident, Google has released Gemini models. cyber security continues to expand its capabilities. In July, the company introduced the Gemini 3.5 Flash Cyber⁠ model, which it developed to find, verify and patch security vulnerabilities. Google built the model on Gemini 3.5 Flash and specifically developed it to quickly find, verify, and fix vulnerabilities. The model can work with more than one agent in CodeMender to examine security problems and prepare a joint report.

Gemini 3.8 Flash Cyber⁠, which arrived in September, is a more advanced version of these works. Google model vulnerability discovery and automatic patching focuses on his duties. In the company’s internal testing covering 20 programming languages, the model discovered a vulnerability over 70 percent success rate reached. In the CWE-Bench test, Gemini 3.8 Flash Cyber ​​failed on the first try. 47.2 percent recorded success rate.

Google also uses Gemini 3.8 Flash Cyber ​​in its own software. The Chrome security team believes the model is better than larger commercial models. 2.6x more correct patches He said he was preparing it. The Google Cloud vulnerability research team also found a critical core vulnerability in less than two hours with the help of the model, which would normally take months. The Chrome security team⁠ previously announced that it uses Gemini-based tools for vulnerability discovery, bug investigation and patching processes.

The company also provides limited access to advanced Gemini cyber defense tools through the Fairwind Program⁠. The program covers select government agencies, Google Cloud customers and security partners. Participants can use CodeMender with Gemini 3.8 Flash Cyber ​​to find and verify vulnerabilities and prepare code fixes. Google initially limits this access to trusted defense teams.

Cyber ​​security studies of the Gemini family started before the 2026 models. Google will launch the Sec-Gemini v1⁠ model in 2025 with threat analysis, root cause analysis and introduced it for tasks aimed at understanding the impact of vulnerabilities. According to the company’s results, Sec-Gemini v1 outperforms other compared models in the CTI-MCQ test. 11 percent left behind. In the CTI-Root Cause Mapping test, the difference is at least 10.5 percent reached the level.

Google uses multiple layers of security to limit the access of web-enabled agents like Gemini. The measures he announced for Chrome⁠ include a separate controller model, access restrictions to web resources, and user approval for critical operations. The company also adds real-time threat detection and security testing to these measures. These methods specifically aim to limit which resources agents can access if they encounter untrusted web content.

The path followed by Gemini in the test in May revealed three concrete security elements. Using internet access, the model found real companies, performed password guesses, and evaluated credentials in public repositories. Gemini did not proceed when it realized it was facing real targets, and Google notified the three organizations. While Irregular troubleshoots the test environment, Google continues to improve Gemini’s cybersecurity models with limited access programs and additional security measures.

Danish Kapoor