Product Launch2 min read

Google Adds Computer Control to Gemini 3.5 Flash

June 24, 2026Synthesized from 2 sources: Google DeepMind, TLDR AI

Google has built the ability for its Gemini 3.5 Flash AI to see and control a computer screen, browser, or mobile device, moving AI agents from answering questions toward doing actual office work on your behalf.

AI tools that answer questions are now giving way to AI agents that take actions. Google's move to embed screen-control into Gemini 3.5 Flash is a clear sign of where this is heading.

The basic idea: instead of a person sitting at a computer, an AI can now see the screen, read what is on it, and take actions as if it had a mouse and keyboard. It can work across a browser, a mobile app, or a desktop application. Google released this as a built-in feature of Gemini 3.5 Flash, the model it already makes available to businesses and developers through its enterprise platform and API.

This matters because Gemini 3.5 Flash is not a niche research product. It is the model Google has rolled out to billions of users through the Gemini app and AI Mode in Search. Adding computer-control to this model means the building block for this kind of automation is now widely accessible, not locked inside a specialist tool.

To understand the competitive picture: Anthropic's Claude has had Windows desktop computer-control since February 2026, letting it open apps, browse the web, and fill in spreadsheets. Google's approach has leaned toward browser and mobile automation, which is where most back-office SaaS tools live. So the practical territory each covers is somewhat different.

The safety question is not a side note. When you give an AI the ability to click and type on your systems, a new class of attack becomes possible. A malicious instruction hidden inside a document the AI reads, or a webpage it visits, can redirect the AI to do something harmful. This is called an indirect prompt injection, and it ranked as the number-one security risk for AI systems in OWASP's 2025 report on AI vulnerabilities. In late 2025, researchers demonstrated a real exploit in an enterprise AI assistant that let attackers compromise protected data with no employee action required.

Google's response includes adversarial training, meaning the model was specifically trained to resist these tricks. It also offers two optional controls for businesses: a confirmation step that asks a human to approve sensitive or irreversible actions before the AI takes them, and an automatic stop if the system detects a suspicious hidden instruction. These are opt-in, not on by default, which means businesses need to actively turn them on.

For any organisation considering this kind of automation, the human-approval safeguard is the one worth enabling from day one. An AI acting on your systems without a checkpoint on high-stakes actions is a governance problem, not just a technical one.

The broader trend here is that the AI market has moved decisively from tools that assist to systems that act. Global investment in AI agents is projected to grow from around $5 billion in 2024 to over $47 billion by 2030. The gap right now is not technology; it is the distance between running a pilot and actually deploying these agents in production. Most organisations experimenting with agents have not yet scaled them to real operations.

For non-technical managers, the near-term question is simple: which repetitive, screen-based tasks in your operation could a well-supervised AI agent handle? Supplier portal checks, compliance documentation reviews, data entry across multiple systems, and routine reporting are all plausible starting points. The technology is ready enough to test. The supervision setup, meaning who reviews what the agent did and how errors get caught, is what most teams have not figured out yet.

Stay informed

Get AI intelligence like this delivered to your inbox.