Webcmd Gives Your Agents a Sitemap Memory So They Stop Rediscovering Sites
If you've ever watched a browser agent burn through tokens figuring out the same login flow for the fifth time, you already know the problem. Every run starts from zero. The agent navigates, fails, backtracks, and relearns a site it's already visited a dozen times before. Webcmd is a bet that agents shouldn't have to keep paying that tax.
What It Does
Webcmd is self-learning browser infrastructure for AI agents. It pairs live browser control with a memory layer that captures what an agent learns about a site as it works, then reuses that knowledge on subsequent runs. The stated goal is straightforward: cut browser-agent token spend by up to 90% by eliminating redundant rediscovery.
The architecture is split into two layers. Layer 0 is live browser control, which you reach for when a site is unfamiliar. Through webcmd browser, an agent can inspect pages, click, type, extract content, capture network calls, and complete a task in a real browser. Layer 1 is sitemap memory, which kicks in once a site is familiar but the full action space isn't known. Here, Webcmd captures an agent-facing sitemap of observed pages, states, actions, workflows, APIs, pitfalls, and fallback paths. That map becomes local memory the agent can consult instead of re-deriving.
It's distributed as an npm package (@agentrhq/webcmd), requires Node.js 20.6 or later, and installs a single skill called webcmd-browser into your chosen harness. The README mentions Claude, Codex, and other supported harnesses, plus a custom skills path option.
Why It's Cool
-
It treats navigation knowledge as a first-class artifact. Most browser automation tooling focuses on the mechanics of clicking and typing. Webcmd's interesting move is capturing the context around those actions—workflows, pitfalls, fallback paths, even observed APIs. That's the stuff agents tend to rediscover expensively, run after run.
-
The two-layer split is a sensible mental model. Live control for the unknown, memory for the known. It maps cleanly onto how you'd actually decide whether to send an agent into a site cold or hand it a playbook. You don't have to guess which mode applies—the README frames it as a scenario-based choice.
-
Profile support for logged-in sessions. The examples show tasks like summarizing unread LinkedIn messages using a
workprofile, or collecting X bookmarks with asocialprofile. That's a practical detail—authenticated browsing is where a lot of real agent work lives, and it's often the messiest part to wire up. -
The skill install is deliberately narrow. You install exactly one skill,
webcmd-browser, and you're told to load or tag it only for live browser work. Setup commands don't need it. That's a small but thoughtful decision that keeps the skill from bleeding into places it doesn't belong. -
The example prompts read like real work. Research across Hacker News and Reddit, pulling publication metadata from PubMed, looking up Grainger part numbers, combining Grainger prices with SAP Ariba purchase-order status. These aren't toy demos—they're the kind of multi-source, read-only lookups that agents are actually being asked to do.
-
The token claim is the whole pitch. Cutting browser-agent spend by up to 90% is a specific, falsifiable number. Whether it holds up in your workload is an empirical question, but it's the right thing to optimize for. Browser agents are token-hungry, and caching navigational context is a plausible lever.
How to Try It
The fastest path is to let your agent do the setup. From the README:
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
If you'd rather do it manually, you'll need Node.js 20.6+:
npm install -g @agentrhq/webcmd
webcmd skills add
When prompted, pick your harness—Claude, Codex, another supported option, or a custom skills path. That installs the webcmd-browser skill. Load or tag it only when you're doing live browser work, then describe the outcome you want. For example:
Use webcmd to research the latest discussions about browser automation across Hacker News and Reddit, then return a concise comparison with source links.
One thing worth noting: the README doesn't spell out the full sitemap memory lifecycle in the truncated portion, so if you want details on how the memory is stored, scoped, or invalidated, you'll want to check the docs at webcmd.dev.
The repo is at github.com/agentrhq/webcmd. There's also a Discord community and an X account if you want to follow along.
Final Thoughts
Webcmd is aimed at developers running browser agents against sites they hit repeatedly—research pipelines, procurement lookups, logged-in inbox triage. If your agents are burning tokens re-learning the same navigation every run, the sitemap memory concept is worth an afternoon of testing. It's early-stage infrastructure, and the README leaves some implementation questions open, but the core idea—stop paying agents to rediscover the web—is the kind of framing that makes you wonder why more browser tooling doesn't work this way already.
Follow @githubprojects for more developer tools and open source projects.