Computer-use models
A computer-use model reads a screen and emits mouse or keyboard actions that an application can execute.
A computer-use model connects visual perception to interface actions. OpenAI’s system card describes a model that interprets screenshots and uses a cursor and keyboard. It works with the buttons, menus and text fields people see on a computer screen.
One action is not a workflow. The application executes the model’s request, captures the changed screen and returns that screenshot. Consequential actions should require user confirmation.
Benchmarks expose the gap between a plausible click and a completed task. OSWorld evaluates real web and desktop apps, file operations and workflows spanning applications. Mind2Web contributes more than 2,000 tasks from 137 real websites. Windows Agent Arena includes more than 150 tasks across representative domains.
These systems remain fallible. The original OSWorld study reported 12.24% success for its best evaluated model versus more than 72.36% for people. WebArena reported 14.41% end-to-end success for its strongest GPT-4 baseline. Prompt injection remains a concern, and consequential actions still need human oversight.
Many digital tasks are exposed through human interfaces rather than a purpose-built API.
Observe the current screen, choose an action, execute it, and inspect the result.
- 1 · observeCapture the current screen and pair it with the user's instruction.
- 2 · decideSelect a mouse or keyboard action from the current state.
- 3 · executeRun that action in the browser or operating-system environment.
- 4 · inspectFeed the resulting screen back to the model before the next action.
The surrounding application executes the model's requests. Consequential actions should require user confirmation.
| Who | What they ask | What it works with |
|---|---|---|
| QA engineer | “Exercise a purchase flow in a browser.” | Screenshots and the visible interface state |
| Office worker | “Move information between a document and a form.” | The open applications and files |
| Accessibility workflow | “Operate software through common pointer and keyboard controls.” | The rendered interface |
| Research evaluator | “Measure whether an agent completed a desktop task.” | The task's final application state |
- A visual agent can interact with a GUI without a custom integration for each site.
- OSWorld supplies execution-based evaluation for tasks across web and desktop applications.
- Mind2Web covers more than 2,000 tasks from 137 real websites across 31 domains.
- Windows Agent Arena includes more than 150 Windows tasks across representative domains.
- Screenshot understanding does not make computer use reliable; the original OSWorld evaluation found large gaps from human performance.
- Prompt injection remains a concern when instructions can appear inside the interface.
- Long web tasks remain difficult; WebArena's reported GPT-4 baseline completed 14.41% end to end.
- The model still needs human oversight for consequential actions and unexpected states.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialOperator System Card, OpenAI · read 28 Sept 2026
- paperOSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments, Xie et al. · read 28 Sept 2026
- paperMind2Web: Towards a Generalist Agent for the Web, Deng et al. · read 28 Sept 2026
- paperWebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al. · read 28 Sept 2026
- paperWindows Agent Arena: Evaluating Multi-Modal OS Agents at Scale, Bonatti et al. · read 28 Sept 2026
- docsComputer use, OpenAI · read 28 Sept 2026