The promise has become a genre of its own: an AI agent that lives on your machine, reads your files, runs your commands, browses, writes, and hands back finished work — without sending your data into anybody’s cloud. In 2026, two open-source projects embody that dream better than the rest, and they crush everything else on popularity: OpenClaw and Hermes Agent.
Both are free, open, and designed to run locally. Both promise autonomy. But they are not alike, and one of the two carries a reputation as a walking hazard. We installed them on the same Windows PC, plugged them into the same models, and put them to work under real conditions. A blunt verdict, with figures, and a winner on the ground that interests us: getting an agent to work alone, locally, over time.
Test results: the verdict in brief
No undisputed winner: it is a close match. Across our 7 criteria the score is 3 to 2 in favour of Hermes Agent, with 2 ties — too thin a gap to call it dominance. Each wins its own ground: OpenClaw on native Windows and the breadth of its ecosystem; Hermes on security, memory and stability. Two criteria end level because the data honestly separates nobody: local agentic access (both have full access to the shell and files) and cost (both burn tokens, and the only real lever — the choice of model — is identical). The right choice depends on your profile, not on a scoreboard.
| Criterion | OpenClaw | Hermes Agent | Edge |
|---|---|---|---|
| Installation & native Windows | Native, dedicated app, no admin rights | Native Windows "experimental", WSL2 advised | OpenClaw |
| Breadth of integrations & ecosystem | ~24 channels, marketplace, ~380,000 ★ | 6 channels, ~199,000 ★ | OpenClaw |
| Local agentic work (shell, files) | Native PowerShell, mature tool surface | 40+ tools, snapshots before edits | Tie |
| Memory & continuous learning | File-based memory + manual skills | Persistent memory + self-generated skills | Hermes |
| Security & execution isolation | 8.8 vulnerability patched, exposed instances, booby-trapped skills | Containers as a boundary, permission by default | Hermes |
| Stability & reliability over time | Breaks on updates, unstable WhatsApp | More stable according to reports, but slow CLI | Hermes |
| Cost control & transparency | "Heartbeat" and swelling context = surprise bill | High fixed overhead per call | Tie |
Le Recul's verdict. Neither is the "autonomous digital employee" the marketing sells: both demand careful configuration, supervision, and real discipline on cost.
The match is a draw, or nearly (3-2, 2 ties). Hermes Agent takes a slight edge on what matters for running an agent over time — stability, isolation, memory — but the gap is not decisive.
OpenClaw keeps strong cards: native Windows with no tinkering, the most integrations, the liveliest ecosystem. There is no universal winner: choose on your profile, not on the GitHub star count.
Two agents, two philosophies
Before comparing, you have to understand that these two tools do not start from the same place. Confusing them means picking the wrong project.
OpenClaw was born out of messaging. Created in late 2025 by Austrian developer Peter Steinberger (who joined OpenAI in early 2026), it changed name five times in two months — an amusing detail that says a lot about the project’s speed. One persistent rumour is worth correcting in passing: Anthropic did ask for a name change (the old “Clawdbot” sounded too close to “Claude”), but that only forced an intermediate step; the name “OpenClaw” is a deliberate choice. Its architecture comes down to one word: the Gateway, a single process running on your machine that links all your messaging apps (WhatsApp, Telegram, Discord, Slack, Signal, iMessage and more) to an agent. You write to it from the app you already use, and it acts. Under the MIT licence, the project shows extraordinary popularity: on the order of 380,000 stars on GitHub in mid-June 2026.
Hermes Agent comes from elsewhere. Published in early 2026 by Nous Research (a recognised name in open-weights AI), it bets everything on an agent that learns. Its promise: memory that persists from one session to the next, and a self-improvement loop — after a complex task, the agent can generate a “skill” (a reusable procedure, with its pitfalls and its verification steps) that it will reload next time instead of thinking from scratch. Also under the MIT licence, it counts roughly 199,000 stars at the same date. Its home ground is the server and self-hosting.
One formula captures the opposition well: OpenClaw wraps an agent around a messaging gateway; Hermes wraps a gateway around an agent that learns. The common ground — and the only one this comparison claims to settle — is their ability to work alone, locally, on Windows. Our ranking of operating-system agents tracks the category as a whole, and our ranking of open models covers the engines you can plug into them.
Installation and first launch on Windows
This is the first real gap, and it weighs heavily for a Windows audience.
OpenClaw owns Windows. Its stated promise is “any OS, any platform”, and it holds: it installs natively, either through a dedicated desktop application (no administrator rights, on recent Windows 10 or Windows 11), or from the command line, or inside WSL2 for maximum Linux compatibility. In practice the first launch is quick; the real friction is not installing the software but wiring up the messaging channels, particularly WhatsApp (QR code scan, sessions that expire).
Hermes Agent pushes you towards Linux. Its official platforms are Linux, macOS and WSL2; native Windows support is explicitly labelled “experimental”. On a Windows machine the recommended route therefore goes through WSL2 — an extra layer that is nothing insurmountable for a developer, but which rules out the non-technical Windows user from the start. Installation happens through a single script.
On machine footprint, both stay reasonable at startup (observed orders of magnitude, not laboratory measurements): a base install on the order of a few hundred MB to a little under a GB, which then grows with history, memory and above all the browsing harness — each embeds a Chromium (around 400 MB) for web tasks. At rest the service uses a few hundred MB of RAM; mid-task with the browser open, expect 1 to 2 GB. Nothing extreme: these tools are hungry for tokens, not RAM. It is the API invoice, not the task manager, that hurts — more on that below.
For anyone who wants to stay fully local (model executed on your own machine through Ollama, vLLM or LM Studio), both allow it, but hardware becomes the constraint. A small 8-billion-parameter model answers chat but stalls on chaining tools; for a reliable agent, aim for a 27 to 36 billion parameter model on a card with at least 24 GB of VRAM. Free in tokens, but not in hardware.
Price, models and real cost
Good news first: both programs are free and open source (MIT licence). No subscription for the software itself. Both are “bring your own key”: you plug in the model of your choice, and that is what costs.
Here are the public prices of the models commonly plugged into these agents (official prices recorded in June 2026, in dollars per million tokens, input / output):
| Model | Input ($/M) | Output ($/M) |
|---|---|---|
| Claude Opus 4.8 | 5.00 | 25.00 |
| Claude Sonnet 4.6 | 3.00 | 15.00 |
| Claude Haiku 4.5 | 1.00 | 5.00 |
| GPT-5.5 | 5.00 | 30.00 |
| Gemini 3.1 Pro | 2.00 | 12.00 |
| Qwen3 Max | 0.78 | 3.90 |
| DeepSeek V4 Flash | 0.14 | 0.28 |
The gap is vertiginous: DeepSeek V4 Flash costs roughly 35 times less on input and 90 times less on output than a top-end model. On an agent that endlessly repeats the same operations, the choice of model decides everything — the same arithmetic we ran on DeepSeek V4.1 Flash against GPT-5.6 Luna, where the advertised token price turned out not to be the bill.
Why the invoice explodes
Because these agents burn tokens, and that is the trap nobody sees coming. The mechanism is documented on the OpenClaw side: at every turn, the agent sends the model its full context (configuration files, session history, memory). A reference discussion on the repository describes API limits being hit twice in the same day after only twenty to thirty short messages, with context swelling each turn and cache reads “in the hundreds of thousands of tokens”. Worse: a periodic “heartbeat” queries the model again even when the agent is doing nothing. Concrete result reported by one user: $250 gone on the first day of a “harmless” test. The most spectacular episode — a bill of roughly $1.3 million in one month — belongs to OpenClaw’s own creator, on R&D funded by OpenAI: that is not a normal user’s invoice, but it illustrates how fast this runs away.
Hermes Agent does not escape the rule, for a different reason: a substantial share of every call is a fixed cost (tool definitions, system instructions) before you have even asked your question. The workaround is the same on both sides: drop down a model tier (switching to Haiku typically divides the bill by five) and exploit the cache, where DeepSeek is unbeatable. Hermes also offers its own subscription gateway, Nous Portal (around $20 a month for the entry tier according to third-party sources), convenient but paid. OpenClaw sends you straight to the providers or to OpenRouter.
What they can do: capabilities and autonomy
On paper, both tick almost every box of a modern agent. The devil is in the detail.
| Capability | OpenClaw | Hermes Agent |
|---|---|---|
| Shell & files | Yes (native PowerShell on Windows) | Yes (40+ tools; snapshots before edits) |
| Browser & web search | Yes | Yes |
| Messaging channels | About twenty | 6 |
| Scheduling (cron) & triggers | Yes (cron, webhooks) | Yes (natural language or cron) |
| Sub-agents / delegation | Yes (multi-agent) | Yes (3 parallel sub-agents by default) |
| MCP (external tools) | Yes | Yes |
| Vision / image / voice | Yes (canvas, voice) | Yes (vision, image generation, speech synthesis) |
| Code editor integration | Native mobile/desktop apps | VS Code, Zed, JetBrains (ACP protocol) |
| Persistent memory | Readable Markdown files | SQLite database (opaque) + self-generated skills |
The difference in autonomy comes down mostly to a default setting: OpenClaw, in “single trusted operator” mode, runs shell commands by default without asking for confirmation. That is powerful — and dangerous. Hermes asks for approval by default. On guardrails, Hermes has a conceptual head start: it can run inside a container (Docker, Singularity) or a cloud sandbox, with the container acting as a genuine security boundary. OpenClaw also has a sandbox and tool policies, but they must be switched on manually.
Memory and learning: Hermes’s bet
This is Hermes Agent’s number one argument, and it deserves a closer look — with one important caveat.
Hermes’s memory is real: facts, preferences and procedures persist from one session to the next, and the agent can generate its own skills to save time on repeated tasks. Nous Research claims a gain on the order of 40% on repetitive tasks once the skills library is filled out — an internal figure, not independently verified, to be taken as a sales argument rather than a measurement.
Above all, two traps await the new user. First, the most advanced self-learning function is off by default: without an explicit configuration step (hermes memory setup), you get basic memory, not the promised learning system. Second, memory is bounded: the official documentation sets strict size limits (around 800 tokens for general memory, 500 for the user profile) and the tool returns an error rather than silently truncating. That is healthier for predictability, but it means consolidating it by hand.
Against that, OpenClaw bets on transparency: its memory is stored in Markdown files you can open, read and correct yourself. Less magical, more controllable. Two philosophies again: Hermes bets on automatic learning of memory, OpenClaw on manual control.
Security: the Achilles heel (especially for OpenClaw)
This is the chapter where the match tips, and where Le Recul cannot stay quiet.
OpenClaw went through a public, documented and severe security crisis. Three facts, all sourced:
- A critical vulnerability, CVE-2026-25253, scored 8.8 out of 10 on the NVD: a single booby-trapped link could leak the Gateway’s access token and lead to one-click remote code execution. It has been patched since version 2026.1.29 — but it illustrates the fragility of the architecture.
- Cisco’s AI security research team called these personal agents a “security nightmare” in January 2026, after finding a top-ranked third-party skill that exfiltrated data to an external server and bypassed the guardrails, without the user noticing.
- Researchers (SecurityScorecard, relayed by The Register) counted tens of thousands of OpenClaw instances exposed on the internet in February 2026, many of them poorly protected.
On top of that, a structural point: OpenClaw’s skills marketplace is not vetted, which makes it a supply-chain attack vector. The project reacted (file-system hardening, skill ratings), and its creator says it plainly: there is no such thing as a “risk-free” agent.
Hermes Agent is not blameless, but its record is lighter. A vulnerability, CVE-2026-7396 (score 5.3, medium), concerns a secondary adapter — it appears on the NVD, but through a third party (VulDB) and the official severity assessment is still pending. Above all, Hermes bets on container isolation and permission by default, which reduces the attack surface. Its specific weakness lies elsewhere: its persistent memory is by nature a surface to protect (malicious content could lodge itself there), and the project carries a governance controversy over the origin of its self-improvement loop — a transparency debate we mention without settling.
The lesson applies to both: an agent with access to your shell, your files and your messaging apps is a target. Never lose sight of that behind the convenience. The same question comes up whenever an agent is handed the keys, as we found comparing a free agent against a paid one.
Day-to-day reliability: what breaks
An autonomous agent is only worth something if it holds up over time. On that point, user reports (specialist forums, tickets opened on the repositories) lean towards Hermes — without either being beyond reproach.
On OpenClaw’s side, the pattern of incidents is real, especially around the gateway and WhatsApp: repeated disconnections, reconnection loops that can trigger a WhatsApp account ban (the connection goes through an unofficial route), an “event loop” freeze that locks up every channel, and regressions on certain updates — including one specific to Windows reported in spring 2026. To be fair, though: several of those tickets were closed without a fix (“not planned”) or marked inactive. The honest reading is not “OpenClaw crashes constantly”, but “OpenClaw moves fast and sometimes breaks, and not every robustness request is a priority”. Another sore point, already mentioned: token cost that runs away if you do not bound the context.
On Hermes’s side, the grievances are less serious but real: a command-line interface judged slow to start and sometimes unstable, a feeling of over-engineering, and the need to switch advanced functions on manually. One user even reported a completely silent memory — but that isolated case, closed by its own author, looks more like a configuration problem than a general failure. The consensus among those who switched is clear: Hermes holds up better over the distance, especially on short contexts.
Three concrete scenarios
- You want an assistant that acts from WhatsApp or Telegram, on a standard Windows PC, with no tinkering. → OpenClaw. Native support and the breadth of integrations make the difference — provided you switch the guardrails on (command approval, sandbox) from the start.
- You want a reliable agent that runs for weeks on a server or workstation, and gets better at your recurring tasks. → Hermes Agent. Stability, container isolation and persistent memory come first. Plan for WSL2 or Linux.
- You want the lowest possible cost. → A technical tie, but the real answer is not the agent: it is the model. An economical model (DeepSeek, Qwen) or a local model on a 24 GB GPU collapses the bill, whichever agent you pick. Our cost / performance ranking compares them continuously.
Strengths and limits of each
OpenClaw — strengths: native Windows with no friction; about twenty messaging channels; huge ecosystem and community; transparent, editable memory; maximum autonomy by default. OpenClaw — limits: documented security crisis (8.8 vulnerability patched, exposed instances, booby-trapped skills); shell execution without confirmation by default; gateway and WhatsApp instability; token cost that runs away; occasional breakage on updates.
Hermes Agent — strengths: persistent memory and self-generated skills; better security hygiene (containers, permission by default); more stable over time; bounded, predictable memory; code editor integration (VS Code, Zed, JetBrains). Hermes Agent — limits: native Windows “experimental” (WSL2 advised); self-learning off by default; half as many integrations; slow command-line interface; weak on pure development tasks against dedicated tools.
What to remember
OpenClaw and Hermes Agent tell two versions of the same dream: an AI that works for you, on your machine, without renting anybody’s cloud. But behind the hype — and the hundreds of thousands of GitHub stars — these are two young, fast-moving and demanding pieces of software, which require an informed user and real discipline.
On the ground that interests us, getting an agent to work alone, locally, over time, the match ends almost level: 3-2 in favour of Hermes Agent, with 2 criteria tied. Hermes takes the edge on security, memory and stability; OpenClaw remains the better choice for anyone who wants native Windows out of the box, the most integrations and the liveliest ecosystem — provided they switch the guardrails on and watch the bill. So there is no universal winner: a real but thin advantage for Hermes, and a choice decided by your use case, not by a score.
The marketing promises an autonomous digital employee. The test says something else: two promising assistants that need supervising the way you supervise a gifted, unpredictable intern. AI makes promises — and we verify.
For more, see our companion comparison on Codex against Cowork, the other AI agent duel on Windows, and line the underlying models up in our free comparator.