Out of the Terminal, Hands on Mouse and Keyboard — GitHub Copilot Starts Operating Screens That Have No API
On October 1, GitHub released "Computer Use" in public preview, letting Copilot directly operate the screens of desktop apps. It opens the door to automating legacy software without an API and to E2E testing, while this piece lays out the risks of misoperation and information leaks and the key point
For a long time, coding agents did their work inside code editors and terminals. On October 1, 2026, GitHub crossed one of those boundaries with a new feature called "Computer Use," now available in public preview. Copilot can now read the screen of a desktop app on your behalf, click buttons, type text, and even scroll and drag. It marks the point where a developer tool is beginning to reach into the automation of legacy software that has no API.
Turning the Screen Itself Into a Surface for Action
Computer Use reads not only the data an app exposes, but also the visual context shown on screen. According to GitHub, Copilot can "read what the app displays and the context on screen, click controls, type and edit text, press keys, scroll and drag, and carry out operations that span multiple apps." The aim is clear: automating "GUI-only" software that has neither an API, a command line, nor MCP. Legacy in-house systems that agents could never touch before are now coming into scope.
Pierce Boggan of GitHub, who first announced the feature, posted on X that "computer use has come to GitHub Copilot," citing examples such as automating expense reports and travel bookings, and speeding up end-to-end testing. Early users report that the work proceeds in the background, without the agent hijacking your mouse, keyboard, or the window you're working in.
Where It Runs and How to Enable It
Here is where it's available and how to turn it on.
- Supported environments: GitHub Copilot CLI, and the GitHub Copilot app for macOS and Windows
- Status: Public preview as of October 1, 2026. Disabled by default
- Enabling it: In the CLI, run
/computer on; in the app, enable it under Settings → Computer Use - macOS prerequisites: Accessibility and Screen Recording permissions must be granted
It reportedly works best when you clearly state the outcome you want, the apps involved, and the constraints you want respected. The more precisely you spell out "in this app, under these conditions, up to this point," the less the agent's actions drift.
The Gatekeeping That Stops It From Acting on Its Own
Since the agent touches the screen directly, a mechanism to prevent it from running wild becomes essential. GitHub has adopted a staged permission design.
- Off by default: It does nothing unless you explicitly enable it
- First-time approval: Each time Copilot tries to operate an app, it asks for approval for that specific app
- Saving auto-approval: You can save an approval to automate it going forward, but the saved approval applies to both the app and the CLI
- Organization-level disabling: Administrators can disable the feature entirely in management settings
GitHub itself advises against choosing "always allow" for apps that handle sensitive data, and guides users to prefer APIs, MCP servers, and terminal commands when predictable results are needed, treating Computer Use as a last resort. It's a design that foregrounds convenience while prompting users to narrow its scope.
What Changes for Development and Business
The business implication is that the range of what can be automated expands from "things with an API" to "everything with a screen." The anticipated use cases include the following.
- Summarizing or standing in for operations on GUI-only legacy desktop apps
- Expense reports, and moving data between software that isn't integrated
- Updating presentation materials
- Speeding up end-to-end testing
On the development side, the implications for test automation look especially significant. A path opens up to validate screens that have been neither API-enabled nor scripted, by instructing the agent in words much as you'd write a spec. Legacy in-house systems that were hard to touch for modifications may start being folded into agent-driven operations.
Behind the Convenience: Operational Mistakes and Information Leaks
At the same time, automation that directly operates the screen carries risks of its own. GitHub's own guidance notes that when the screen changes — due to app version differences, for example — "Copilot may click the wrong control, enter text in the wrong field, or repeat an action." There is also the concern that unexpected screen content could inadvertently trigger operations against financial or core business systems. Furthermore, on Windows, the possibility that other people's information could appear on screen is flagged as a point of concern for handling personal data.
Because it presumes broad permissions such as screen recording and accessibility, the range an agent can touch certainly grows wider than before. The gates of off-by-default, per-use approval, and organization-level blocking are the safeguards for that, but if "always allow" is extended carelessly, the gates become gates in name only. If you adopt it, practical boundaries would be to exclude apps handling sensitive data from auto-approval and to start with low-impact tasks. Convenience and the breadth of operating permissions are two sides of the same coin, and the judgment of how much to delegate is left to the user.
References: GitHub Changelog: GitHub Copilot can now interact with desktop apps with computer use / A summary of Pierce Boggan's announcement (X) / Aivy: GitHub Copilot can now use the apps on your computer / Technobezz: GitHub Copilot Gains Computer Use for Desktop Apps


