The dream is the same on both sides: you hand a task to an AI, it opens your files, runs the commands, writes, fixes, and hands back finished work. No more copy-paste, no more manual steering — a digital colleague who works while you do something else.

OpenAI sells that dream under the name Codex. Anthropic sells it under the name Cowork. Both can act on their own on a machine. We installed them on the same Windows 11 PC and put them to work, under real conditions. A blunt verdict, with figures, and a winner on the ground that interests us: agentic work running locally, on Windows.


Test results: the verdict in brief

Winner on the ground tested (autonomous agent running locally on Windows): OpenAI Codex, by a narrow margin — 4 criteria out of 7. Codex wins because it runs natively, installs without friction on a standard Windows machine, and pushes autonomy further. Cowork is not beaten everywhere: it wins on office versatility, on guardrails, and on isolation inside a virtual machine.

Criterion Codex (OpenAI) Cowork (Anthropic) Edge
Windows support & installation Native, recent editions Hyper-V VM, Windows 11 Pro required Codex
Local agentic work (shell, files) Native PowerShell, direct access Through an isolated VM Codex
Versatility beyond code (office work) Limited (code-oriented) Apps, spreadsheets, browser Cowork
Autonomy & background work Parallel threads, cloud tasks No background work, app closed Codex
Guardrails & control Sandbox levels Plan approved before acting Cowork
Entry price & accessibility Free tier + Go at $8 Included from Pro at $20 Codex
Isolation / execution safety System-level sandbox (past vulnerability) Strong VM isolation Cowork

Le Recul's verdict. Neither is the autonomous employee the marketing promises: both require approval and supervision.

To put an AI to work locally on Windows, Codex wins — native, frictionless, more autonomous, and available all the way down to a free tier.

Cowork takes the lead back the moment you step outside code: desk tasks (spreadsheets, browser, documents), stricter guardrails and execution isolated inside a virtual machine. The score is close (4-3): this is a choice of use case, not a show of force.


Two agents, two philosophies

Before comparing, you have to understand that Codex and Cowork are not aimed at the same user. Confusing them means picking the wrong tool.

Codex (OpenAI) was born on the developer side. Its first building block, an open-source command-line interface, appeared in April 2025; a web version followed, then a model dedicated to agentic coding in autumn 2025, and finally a native Windows application with its own sandbox in spring 2026. Today Codex lives everywhere a developer works: the terminal, the code editor, the web, GitHub and the Windows desktop. Its core business: reading, writing, running and fixing code autonomously.

Cowork (Anthropic) targets office work in the broad sense. Launched first on macOS in early 2026, it reached Windows a few weeks later, then became generally available in the spring. Its promise: “delegate a task, get a finished deliverable”. In practice, Cowork opens applications, fills spreadsheets, browses in your browser and handles your files, running by default on a Claude Opus model with a very large context window (on the order of a million tokens).

In other words: on one side a specialist’s tool built for code, on the other a general-purpose assistant built for office work. The common ground — and the only one this comparison claims to settle — is their ability to work on their own, locally, on Windows. Our ranking of operating-system agents tracks this category as a whole.


Price, models and how the subscription works

This is often where the promise meets the invoice. An overview of the consumer plans:

Criterion Codex (OpenAI) Cowork (Anthropic)
Free Yes (Free plan) No — excluded from the free plan
Paid entry point Go: $8/month — Plus: $20/month Claude Pro: $20/month ($17/month billed annually)
"Pro" plan Pro: from $100/month (5× and 20× tiers) Claude Max: from $100/month (5× and 20×)
Default AI model GPT-5.5 / GPT-5.4 / GPT-5.4 mini Claude Opus 4.6 (1M token context)
Billing logic Quotas per 5-hour window, then credits/tokens Included in the subscription, draws on its quota

Codex: five rolling hours, then tokens

Codex does not bill per message. Usage (local messages plus cloud tasks) is counted on a rolling five-hour window. Here is the number of local messages per window, by plan and model:

Plan GPT-5.5 GPT-5.4 GPT-5.4 mini
Plus ($20) 15 to 80 20 to 100 60 to 350
Pro 5× ($100+) 75 to 400 100 to 500 300 to 1,750
Pro 20× ($100+) 300 to 1,600 400 to 2,000 1,200 to 7,000

The Business plan is pay-as-you-go (quotas aligned on Plus) and Enterprise/Edu is quoted on request. Beyond the quotas, or through an API key, Codex switches to billing in credits per million tokens:

Model Input Cached input Output
GPT-5.5 125 12.5 750
GPT-5.4 62.5 6.25 375
GPT-5.4 mini 18.75 1.875 113

OpenAI puts the real cost at roughly $100 to $200 per developer per month, with wide variance depending on the model, the number of agents launched in parallel and use of “fast mode”. The logic is plain: the more the agent works, the higher the bill. One example says it all — launching three parallel agents on GPT-5.5 to rewrite a module drains a Plus quota in a handful of hours, where the same work on GPT-5.4 mini lasts longer but with a less capable model.

Cowork: included in Claude, but it eats your quota

Cowork has no price of its own: it is included in paid Claude subscriptions. The grid:

Claude plan Price Cowork included?
Free $0 No
Pro $20/month — or $17/month billed annually ($200 upfront) Yes
Max from $100/month (5× and 20×) Yes
Team Standard $20/seat (annual) — $25 monthly Yes
Team Premium $100/seat (annual) — $125 monthly Yes
Enterprise per-seat price + usage at API rates Yes

The catch: Cowork consumes the same quota as your Claude plan. A complex session can absorb the equivalent of several dozen chat messages — on the Opus model, credits go fast. “Included” therefore does not mean “unlimited”: delegating a large office task in the morning can seriously eat into what is left for the rest of the day.

Reading the figures. The entry prices look alike (free / around $20 / $100+), but the mechanics differ completely. Codex counts in five-hour windows then bills per token: transparent, but unpredictable once you multiply agents. Cowork is folded into a subscription many people already pay for, but it draws on a shared quota and offers no free tier. On entry price and accessibility alone, Codex’s free tier takes the point — it is the only one of the two you can try without reaching for a card.


What they can do on their own

Capability Codex Cowork
Read / write local files Yes (working folder) Yes (authorised folders)
Run shell commands Yes (core business) Through the VM and its tools
Edit multi-file code Yes (IDE + GitHub) Possible, not its main axis
Drive apps / browser / spreadsheets Limited (code-oriented) Yes (main axis)
External connectors (MCP) Yes + web search Yes (plugins + connectors)
Parallel agents Yes (parallel threads) Not promoted
Control from mobile Through the ChatGPT iOS app Yes

The picture in use is clear: Codex covers development in depth (running commands, multi-file editing, GitHub integration, parallelism), Cowork covers office work in breadth (applications, spreadsheets, browser). The two scopes barely overlap, which is why the “match” only makes sense on the common ground: acting on the machine, locally. To drive the system, run commands and touch the file system — the very definition of local agentic work — Codex is more direct; to turn a pile of documents into a deliverable, Cowork is more at ease.


Autonomy and guardrails

This is where the two cultures diverge most.

Codex grades autonomy through sandbox levels:

  • read-only: it inspects, touches nothing without approval;
  • write inside the working folder (the default mode): it reads, edits and runs commands inside the project, but asks permission before going online or outside the folder;
  • full access: no boundaries left, neither files nor network.

On Windows, that sandbox leans on system mechanisms: depending on configuration, Codex isolates its actions through reduced-privilege accounts, file-system boundaries and firewall rules, or through a restricted token and access control lists. The direct consequence: Codex can be set to act very far without human intervention — including running tasks in the cloud or running several agents at once. That is its strength for autonomous work… and exactly what demands discipline.

Cowork bets on the approved plan. Before acting, it shows what it intends to do and waits for your go-ahead; you choose the folders and connectors it can reach, and you can redirect it at any step. That systematic control is reassuring — it limits nasty surprises — but it breaks the idea of an agent that gets through the work while you sleep. Above all, Cowork does not run in the background once the application is closed: tasks only run if the computer is on and the app is open.

Criterion verdict: on raw autonomy (parallelism, background, cloud offload), Codex takes the point. On guardrails (approved plan, step-by-step control), Cowork takes its own. Two opposite visions of what “an agent that works on its own” should be. We applied the same test to autonomous agents at the free end of the market in our comparison of a free agent against a paid one.


Windows locally: native against virtual machine

This is the heart of the matter, and it is at installation that the difference jumps out.

On Windows Codex Cowork
Execution mode Native (app, terminal, IDE extension) Hyper-V virtual machine
Shell Native PowerShell, no WSL required Goes through the isolated VM
Isolation System-level sandbox Strong VM isolation
Windows edition required Windows 11 recommended (recent Win 10 tolerated) Windows 11 Pro/Enterprise (Home insufficient)
Installation friction Low High (depends on the VM)
Background tasks Yes (cloud) No if the app is closed

Codex plays the native card. On our machine it ran as close to the system as possible: native PowerShell commands, a sandbox managed directly by Windows, no dependency on WSL. You hand it a task, it reads the folder, runs, fixes — the experience is immediate, on a standard Windows edition.

Cowork plays the virtual machine card. Cowork installs a Hyper-V virtual machine service on the PC and runs its work inside it. That is elegant on the security side — the agent is walled off from the rest of the machine — but it imposes a brutal prerequisite: Windows 11 Pro or Enterprise. The Home edition does not provide the virtual-machine management component required, and installation stops before it starts. Add the hiccups we saw on Windows: a VM service that dies after sleep, network conflicts with a VPN or Docker that leave the VM without internet access. When the VM coughs, the agent loses its hands.

On this decisive criterion — putting an agent to work locally on an ordinary Windows machine — the gap is clear. Codex installs and works anywhere; Cowork demands a Pro edition and a virtualisation layer that can get in the way. It is one of the reasons the overall balance tips towards Codex for local use on Windows. That said, this awkward VM protects the rest of your system better: the price of convenience is less isolation.


Reliability, security and hidden costs

An agent that acts on its own on your machine is also a risk that acts on its own on your machine. Both tools have blind spots, and it would be dishonest to spare either.

Codex — security already caught out. A critical vulnerability (CVE-2025-61260) targeted Codex CLI: a booby-trapped repository could, through a simple local configuration file pointing at malicious definitions, run a command as soon as the tool started, without any consent. Enough to install remote access, steal credentials, exfiltrate secrets or compromise a build pipeline. The flaw was patched in a later version, but it is a reminder of a golden rule: an agent that reads the configuration of an unknown project should be treated as code you are executing. On top of that, the cost climbs mechanically with use, because of per-token billing.

Cowork — deletion risk and dependence on the service. The danger is not an exotic vulnerability, it is the mundane: several users report damage after an instruction that was too vague — typically an agent asked to “tidy up a folder” that erases far more than expected. Cowork also depends on a remote service: during congestion it inherits the slowdowns and outages of the Claude service, and it does not work offline or with the application closed. Its security asset remains VM isolation, which limits the damage when things go wrong.

Our reading: a draw on intent (both can do damage), but different risk profiles. Codex mainly exposes you to uncontrolled execution; Cowork mainly to data deletion and service outage.


Three concrete scenarios

To get past generalities, here is how each tool fares on three typical tasks for a Windows workstation.

1. Fixing a bug in a codebase. Reading the repository, editing several files, running the tests, fixing in a loop. This is Codex’s playground: direct shell access, multi-file editing, Git integration. Cowork can get close, but without the same fluency. Edge: Codex. For the models behind this kind of work, see our AI ranking for code and our comparison of the best AI for coding.

2. Compiling a report from Excel files and a website. Opening several spreadsheets, extracting figures, cross-referencing with a web page, producing a clean document. Here Cowork is in its element: it drives applications and the browser, and its large context helps it keep track. Codex, code-oriented, is less natural at it. Edge: Cowork.

3. Launching a long task and letting it run. A codebase clean-up, a migration, a batch job you want to start in the evening. Codex can offload the task to the cloud and run agents in parallel; Cowork stops the moment the application is closed. Edge: Codex.

Two scenarios out of three go to Codex — not because it would be “better” in the absolute, but because two of those three tasks are local system work, its strong suit. On the pure office task, Cowork takes the lead without argument.


Strengths and limits of each

Tool Strengths Limits
Codex Native on Windows (PowerShell); deep on code; parallel threads; background cloud tasks; free tier. Autonomy that can go all the way to full access; past security vulnerability (CVE); per-token billing that climbs; little use outside development.
Cowork General-purpose (apps, spreadsheets, browser); plan approved before acting; VM isolation; included in the Claude subscription. Requires Windows 11 Pro (Hyper-V VM); no background work with the app closed; shared quota that empties fast; deletion risk; depends on the remote service.

What to remember

  • Winner on local agentic work on Windows: Codex (4-3). Native, no installation friction, more autonomous, available down to a free tier.
  • Cowork wins outside code: office work (spreadsheets, browser, documents), stricter guardrails, isolation inside a virtual machine.
  • Entry prices are close (free / around $20 / $100+), but the mechanics are opposite: Codex counts in five-hour windows then bills per token (around $100-200 per developer per month according to OpenAI); Cowork is included in the Claude subscription but draws on the shared quota.
  • Autonomy: Codex can be pushed to full access and run in the background; Cowork imposes an approved plan and stops once the app is closed.
  • Security: Codex had a critical vulnerability, now patched (CVE-2025-61260); Cowork mainly exposes you to deletion risk and depends on a remote service.
  • The prerequisite that decides it: Cowork requires Windows 11 Pro/Enterprise; Codex runs on a standard edition.

For the neighbouring duel, this time between two open-source agents that also run locally on Windows, see our comparison of OpenClaw against Hermes Agent.

Our verdict. AI promises a colleague who works entirely on their own; Le Recul checks: what we have are supervised assistants, not autonomous employees — for both of them. On the precise ground of this test — putting an agent to work locally, on Windows — Codex wins, because it installs anywhere, acts closest to the system and pushes autonomy further. Cowork keeps a real lead the moment delegated office work is involved, with stricter guardrails. The close score (4-3) says the essential: this is not a show of force, it is a choice of use case. Either way, the rule stands: back up your data, approve the plans, and never hand full powers to a machine on the strength of a promise.

You can line up the models behind these agents in our free comparator, which recalculates on every ranking update.