TrendWhat rose yesterday, every day at 07:30 KST · 한국어

GitHub · as of October 8, 2026 · View on GitHub →

microsoft/CUAWright

Browser and desktop agent framework

Automate browser and desktop tasks by giving coding models a terminal

The model writes code to handle repetitive tasks such as collecting information from websites or changing settings in desktop apps. Unlike previous methods that predicted the next click by observing the screen state, this tool follows a flow where the model writes and executes scripts in a terminal and checks the results. Code and logs in the local workspace become the core state instead of browser sessions, allowing for more stable execution of complex multi-step tasks.

Browser and desktop integration

On October 3, 2026, the project was renamed from Webwright to CUAWright, integrating support for desktop tasks in an Ubuntu virtual machine in addition to browser automation. Existing webwright commands and Python imports continue to work, but browser tasks are executed with the cuawright web command and desktop tasks with cuawright desktop. This update enabled the reproduction of the OSWorld-V2 benchmark, and the browser runtime moved to the cuawright.webwright path. These structural changes reflect a direction to minimize and integrate the digital agent interface.

Controlling browsers and desktops with code

The tool provides a terminal to the coding model, allowing it to directly manipulate browser or desktop applications. The model writes Playwright scripts to query elements and wait for conditions, handling dynamic behaviors such as lazy loading or re-rendering. Generated scripts can be re-executed or modified for other tasks, reducing the inefficiency of exploring from scratch every time. For desktop tasks, the model inspects applications and executes commands via the terminal in an OSWorld-V2 Ubuntu VM, then checks screenshots to complete the task.

Benchmark performance and cost

Compared to results from running the same model in other harnesses, CUAWright recorded higher scores on all benchmarks. On OSWorld-V2, it achieved a 67.9% partial score using GPT-5.6 Sol, which is 5.2%p higher than the model running alone. For CAD tasks, it recorded an average IoU of 79.0% on BenchCAD and a composite score of 55.7% on CADGenBench. Specifically, when using GPT-5.6 Sol on OSWorld-V2, it reduced the API cost per task by $9.5 while improving the score.

Converting scripts into skills

Skill Factory is an optional extension that refines scripts generated during the solving process into validated parameterized code and stores them in a library. Learned skills run independently without a model in about 40 seconds with zero token cost. Applying this reuse feature to the WebArena benchmark increased holdout accuracy from 55% to 70%, a 15%p rise. It works as a plugin without modifying the agent loop, checking the library before starting a new task to either execute matching skills directly or inject them as hints into the prompt.

Pre-execution checks

Browser tasks require Python 3.10 or higher, Chromium installed via Playwright, and an API key from OpenAI, Anthropic, or OpenRouter. Desktop tasks require Python 3.12 or higher, a Responses-compatible model API key, Docker, and access to official task and asset snapshots that require permission from Hugging Face. Installation involves running pip install -e . and playwright install chromium, while desktop features are added with pip install -e ".[desktop]". API keys must be stored in user-owned files outside the source, task, asset, and result directories with 0600 permissions. Existing webwright commands maintain backward compatibility, but new features require the cuawright command.

By the numbers

Language
Python
Last commit
October 6, 2026

Written by AI from this repository's README on October 7, 2026. GitHub's original is the reference.

View on GitHub →