Back to home@ShadowMiner

hermes-web-mcp

No description

Stars
0
Language
JavaScript
Created
Sep 8, 2026
Updated
Sep 8, 2026
GitHub repo

Introduction

hermes-web-mcp

Web powers for any MCP client — no API keys required.

hermes-web-mcp is an MCP (Model Context Protocol) server that gives DSH Desktop and any MCP client a complete web toolkit:

  • web_search — keyless search via cn.bing direct, with a DuckDuckGo-over-proxy fallback; no search API key needed.
  • web_extract — scrape any page with full JS rendering (local playwright-service + html-to-md chain over dual CDP Chrome) and return clean Markdown.
  • 7 × page_* — real browser automation on a shared CDP Chrome inside an isolated context (never touches your host browser tabs): page_open, page_click, page_type, page_read, page_shot, page_back, page_close.

Tools

ToolPurposeKey arguments
web_searchKeyless web search (cn.bing direct, DDG fallback) → title/URL/snippet listquery (string, required), max_results (number, default 8)
web_extractRender + scrape a page to Markdown via playwright-service + html-to-md; dual-CDP failover, then plain-fetch degradeurl (required), via_proxy (boolean), wait_ms (number, default 2500)
page_openOpen a URL in the shared CDP Chrome; returns the first 4000 chars of body texturl (required), via_proxy (boolean), wait_ms (default 3000)
page_clickClick an element matched by a CSS selectorselector, wait_ms (default 1200)
page_typeFill an input field (optionally press Enter)selector, text, enter (boolean)
page_readRead up to 12000 chars of the current page body
page_shotScreenshot the current page into shots/ and return the file pathfull_page (boolean)
page_backNavigate back one page
page_closeClose the isolated browser context (cleanup)

9 tools in total: web_search + web_extract + 7 page_* tools.

How it works

┌──────────────┐   stdio (MCP)   ┌────────────────────── hermes-web-mcp ──────────────────────┐
│ MCP client   │ ───────────────▶ │ web_search   → cn.bing (keyless) / DDG fallback           │
│ (DSH, etc.)  │                  │ web_extract  → playwright-service (:3003) → html-to-md    │
└──────────────┘                  │                (:8080)  ⇄ CDP Chrome 9333 (direct) /       │
                                  │                9334 (via HTTP proxy)                       │
                                  │ page_*       → playwright chromium.connectOverCDP          │
                                  └───────────────────────────────────────────────────────────┘

web_extract never needs a scraper API key: a local playwright-service renders the page over CDP, and an html-to-md service converts the HTML to Markdown. Two Chrome instances back it:

  • CDP direct (default, 127.0.0.1:9333) — for domestic sites and pages where your normal login session lives.
  • CDP proxied (127.0.0.1:9334) — Chrome with traffic routed through your local HTTP proxy, for overseas/restricted sites.

web_extract tries your preferred CDP first and fails over to the other automatically; if both fail it degrades to a plain fetch (no JS rendering). The page_* tools use Playwright's connectOverCDP against the same shared Chrome in an isolated context — host tabs are never touched.

Environment variables

VariableDefaultMeaning
WEBMCP_PWD:/path/to/playwright-service/node_modules/playwrightPath to a Playwright install whose chromium driver is used for connectOverCDP
WEBMCP_SCRAPEhttp://127.0.0.1:3003/scrapeplaywright-service scrape endpoint (POST {url, wait_after_load, cdp_url}{content, pageStatusCode, contentType})
WEBMCP_CONVERThttp://127.0.0.1:8080/converthtml-to-md endpoint (POST {html}{markdown})
WEBMCP_CDP_DIRECThttp://127.0.0.1:9333CDP endpoint of the direct Chrome
WEBMCP_CDP_PROXYhttp://127.0.0.1:9334CDP endpoint of the proxied Chrome

All variables are optional — omit them and the defaults above are used.

Prerequisites (deployment)

  1. Node.js ≥ 22, then npm install in this repo (only runtime dependency: @modelcontextprotocol/sdk).
  2. playwright-service — a local service exposing POST /scrape {url, wait_after_load, cdp_url}{content, pageStatusCode, contentType}, rendering the page with Playwright against the target CDP. (The firecrawl-lite project has a reference implementation.)
  3. html-to-md — a local service exposing POST /convert {html}{markdown}.
  4. Two CDP Chrome instances:
    • chrome --remote-debugging-port=9333 (direct)
    • chrome --remote-debugging-port=9334 launched with traffic routed through your local HTTP proxy (for sites not directly reachable).
  5. A Playwright install resolvable at WEBMCP_PW (or the default path) for the page_* tools.

The defaults assume services on 127.0.0.1:3003 / 127.0.0.1:8080 and Chrome on 9333 / 9334. Different setup? Just set the env vars.

Registering the server

Generic MCP config (e.g. Claude Desktop claude_desktop_config.json):

{
  "mcpServers": {
    "hermes-web": {
      "command": "node",
      "args": ["D:/path/to/hermes-web-mcp/hermes-web-mcp.js"],
      "env": {
        "WEBMCP_PW": "D:/path/to/playwright-service/node_modules/playwright",
        "WEBMCP_SCRAPE": "http://127.0.0.1:3003/scrape",
        "WEBMCP_CONVERT": "http://127.0.0.1:8080/convert",
        "WEBMCP_CDP_DIRECT": "http://127.0.0.1:9333",
        "WEBMCP_CDP_PROXY": "http://127.0.0.1:9334"
      }
    }
  }
}

dsh-mcp-client (SQLite registry, insert example — adapt column names to your schema):

INSERT INTO mcp_servers (name, command, args, env) VALUES (
  'hermes-web',
  'node',
  JSON_ARRAY('D:/path/to/hermes-web-mcp/hermes-web-mcp.js'),
  JSON_OBJECT(
    'WEBMCP_PW',         'D:/path/to/playwright-service/node_modules/playwright',
    'WEBMCP_SCRAPE',     'http://127.0.0.1:3003/scrape',
    'WEBMCP_CONVERT',    'http://127.0.0.1:8080/convert',
    'WEBMCP_CDP_DIRECT', 'http://127.0.0.1:9333',
    'WEBMCP_CDP_PROXY',  'http://127.0.0.1:9334'
  )
);

Smoke test

With all prerequisites running:

npm install
node smoke.js

smoke.js spawns the server over real MCP stdio, lists the registered tools, then calls web_search, web_extract, page_open (proxied) and page_shot, printing each output. It requires the playwright-service, the html-to-md service and both CDP Chrome instances to be up.

License

MIT — Copyright (c) 2026 ShadowMiner.