dsh-computer-use
DeepSeek Harness desktop-control plugin (dsh-computer-use): virtual mouse/keyboard, screen capture, and driver-level input
- Stars
- 0
- Language
- JavaScript
- Created
- Sep 12, 2026
- Updated
- Sep 12, 2026
Introduction
dsh-computer-use
A computer-use plugin for DeepSeek Harness, adapted from the OpenAI Codex computer-use action vocabulary: the agent "sees" the real Windows desktop through screenshots and "operates" it through a virtual mouse. Each action is delivered directly to the window at the target coordinates, so the user's physical mouse and keyboard are completely unaffected — the user can keep using their own computer while the agent works (Codex virtual-mouse semantics). Zero npm dependencies, structurally identical to dsh-browser-use / dsh-wallpaper.
Tool surface (7 tools)
| Tool | Purpose |
|---|---|
computer_screenshot | Captures the full desktop (multi-monitor), JPEG downscaled to ≤2000px, returned as a model-visible image |
computer_click | Click / double-click / triple-click with left/right/middle button |
computer_move | Moves the virtual mouse without clicking (hover events reach the window below) |
computer_drag | Drags along a path |
computer_type | Types text into the window at the virtual cursor (supports Chinese, bypasses IME) |
computer_key | Presses a key/combination into the window at the virtual cursor (ctrl+a, alt+f4, win, ...) |
computer_scroll | Vertical/horizontal scroll (Playwright semantics: positive y = down, positive x = right) |
Work loop
Codex's computer-use loop: screenshot → model reads the image and decides → action → screenshot to verify. Coordinates are pixel coordinates in the screenshot image (origin at top-left); the tool internally converts them back to real screen coordinates using the scale factor of the most recent screenshot (the attachment store caps a single image at ≤3.5MB / longest edge ≤2000px, so a 4K screen is downsampled, and the coordinate back-mapping is done by the driver layer). Every action returns the cursor position (in screenshot coordinates) and the foreground window (process name + title), so the agent can confirm focus before typing.
Implementation notes
- Virtual input (default):
win-vinput.ps1usesPostMessageto deliver mouse/keyboard messages directly to the window at the action coordinates (WindowFromPoint+ChildWindowFromPointExrecursively drills down to the top-most control; coordinates are converted to client-area coordinates). The physical mouse/keyboard are never touched — user and agent each use their own. The driver layer tracks the virtual cursor position: actions without coordinates (type/key) land at the current virtual cursor position (screen center before the first action); the system prompt tells the agent to "click the input box first, then type". - Physical input (optional): with
inputMode: real,win-input.ps1injects viaSendInput, which really moves the physical cursor; UIPI interception (when the foreground is an elevated window) reports an explicit error. - Screenshot: GDI+
CopyFromScreenof the whole virtual desktop (VirtualScreen, including negative-coordinate secondary monitors);SetProcessDPIAwareensures true pixels. JPEG uses the GDI+ default quality; when over 3.5MB it downscales step by step. - Script split: the three driver scripts are split into
win-input.ps1(physical input) /win-vinput.ps1(virtual input) /win-shot.ps1(screenshot) — merging input injection and screenshots into a single script gets blocked by Windows Defender AMSI (ScriptContainedMaliciousContent). Likewise, the screenshot script cannot use the JPEG quality encoder (EncoderParameters); it must save with default quality only. - Screenshot into the model: goes through the attachment service
(
saveImage) → render returns animageblock, the same pipeline asread_image; the screenshot tool is gated by model capability (models that do not declare image input are rejected outright). - Safety: actions are serialized (
isConcurrencySafe: false); in virtual mode the process blacklist is rejected on the driver side before delivery (the JS side adds a second check); the system prompt requires the agent to confirm destructive actions first. - Driver scripts are pure ASCII (PowerShell 5.1 parses UTF-8-without-BOM as GBK).
Configuration (desktop.yml)
- id: computer-use
name: '@deepseek-ai/dsh-computer-use'
config:
inputMode: virtual # virtual (default, virtual mouse) / real (physical injection)
powershellPath: '' # defaults to powershell.exe from PATH
maxImageDimension: 2000 # screenshot longest edge in pixels (do not exceed the attachment cap)
blockedProcesses: [] # target process blacklist (lowercase, without .exe), e.g. ['cmd','regedit']
Install / uninstall
install.ps1 / uninstall.ps1 (the same three steps as dsh-wallpaper: copy the
package → farm junction → data-dir overlay). After installing, restart DeepSeek
Harness and tell the agent "take a screenshot of my screen".
Known limitations
- Windows only (PowerShell 5.1+ / .NET GDI+).
- Dynamic screen content (video/games) is captured as a static frame; that is expected behavior.
- Virtual mode delivers synthetic window messages: a few applications ignore them (games, raw-input programs, some global hotkeys); the system prompt already requires the agent to report honestly rather than blindly retry when an action has no effect.
- In virtual mode, dialog keys such as Enter are routed according to the application's own focus semantics (e.g. activating the button that has focus), consistent with physical key behavior; that is expected behavior.
- In physical mode (
real), input on elevated windows (UAC high integrity) is blocked by UIPI and reported as an error. - Cannot inject on the lock screen / secure desktop (login screen).
Verification
Four scripts under test\ (all runnable directly, no API key required):
verify-load.mjs: plugin load validation (4 checks) — registers all 7 tools using the real dsh-tools compiler from the runtime, specifically guarding against "schema violations are only discovered after install, killing the backend".verify-driver.mjs: screenshot / cursor / coordinate mapping / physical key basic path (7 checks).verify-vinput.mjs: virtual input end-to-end (14 checks) — real listening windows receive messages, exact coordinates, Enter dialog-key routing, the physical cursor is never moved, and the blacklist rejects before delivery.verify-amsi.mjs: all three driver scripts pass Windows Defender AMSI.
Schema hard constraint (from the 2026-08-24 backend-startup incident, guarded
by verify-load.mjs): dsh-tools' value-schema DSL requires every type: object
node (at any depth) to explicitly declare the boolean additionalProperties; in
the parameter table, required may only be true or omitted (optional parameters
must not write required: false). A violation throws UNSUPPORTED_SCHEMA at tool
registration and kills the backend process.