Concepts

Computer-use models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A computer-use model reads a screen and emits mouse or keyboard actions that an application can execute.

1 · What it is

A computer-use model connects visual perception to interface actions. OpenAI’s system card describes a model that interprets screenshots and uses a cursor and keyboard. It works with the buttons, menus and text fields people see on a computer screen.

One action is not a workflow. The application executes the model’s request, captures the changed screen and returns that screenshot. Consequential actions should require user confirmation.

Benchmarks expose the gap between a plausible click and a completed task. OSWorld evaluates real web and desktop apps, file operations and workflows spanning applications. Mind2Web contributes more than 2,000 tasks from 137 real websites. Windows Agent Arena includes more than 150 tasks across representative domains.

These systems remain fallible. The original OSWorld study reported 12.24% success for its best evaluated model versus more than 72.36% for people. WebArena reported 14.41% end-to-end success for its strongest GPT-4 baseline. Prompt injection remains a concern, and consequential actions still need human oversight.

2 · Why it exists

Many digital tasks are exposed through human interfaces rather than a purpose-built API.

Visual interfaceButtons, menus and text fields can be interpreted from screenshots.
Common controlsMouse and keyboard actions can operate an existing graphical interface.
Long workflowReal tasks may cross web pages, desktop applications and files.
3 · How it works

Observe the current screen, choose an action, execute it, and inspect the result.

Computer use is a closed interaction loop, not a single prediction.
  1. 1 · observeCapture the current screen and pair it with the user's instruction.
  2. 2 · decideSelect a mouse or keyboard action from the current state.
  3. 3 · executeRun that action in the browser or operating-system environment.
  4. 4 · inspectFeed the resulting screen back to the model before the next action.

The surrounding application executes the model's requests. Consequential actions should require user confirmation.

4 · Where it's used
WhoWhat they askWhat it works with
QA engineer“Exercise a purchase flow in a browser.”Screenshots and the visible interface state
Office worker“Move information between a document and a form.”The open applications and files
Accessibility workflow“Operate software through common pointer and keyboard controls.”The rendered interface
Research evaluator“Measure whether an agent completed a desktop task.”The task's final application state
5 · What it solves, and what it doesn't
solves
  • A visual agent can interact with a GUI without a custom integration for each site.
  • OSWorld supplies execution-based evaluation for tasks across web and desktop applications.
  • Mind2Web covers more than 2,000 tasks from 137 real websites across 31 domains.
  • Windows Agent Arena includes more than 150 Windows tasks across representative domains.
doesn't solve
  • Screenshot understanding does not make computer use reliable; the original OSWorld evaluation found large gaps from human performance.
  • Prompt injection remains a concern when instructions can appear inside the interface.
  • Long web tasks remain difficult; WebArena's reported GPT-4 baseline completed 14.41% end to end.
  • The model still needs human oversight for consequential actions and unexpected states.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialOperator System Card, OpenAI · read 28 Sept 2026
  2. paperOSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments, Xie et al. · read 28 Sept 2026
  3. paperMind2Web: Towards a Generalist Agent for the Web, Deng et al. · read 28 Sept 2026
  4. paperWebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al. · read 28 Sept 2026
  5. paperWindows Agent Arena: Evaluating Multi-Modal OS Agents at Scale, Bonatti et al. · read 28 Sept 2026
  6. docsComputer use, OpenAI · read 28 Sept 2026