← Back to home@GooDAnDReaDY

dsh-voice

Voice input for DeepSeek Harness: dictation and voice messages with multi-provider STT fallback chains

Stars
6
Language
—
Created
Aug 19, 2026
Updated
Sep 19, 2026

Introduction

📦 @goodandready/dsh-voice

Zero-Latency Streaming Dictation & Multi-Provider Voice Input for DeepSeek Harness

npm version license DSH Plugin Node version

All Author Projects

🇬🇧 English • 🇷🇺 Русский • 🇨🇳 中文说明

⭐ If you like this plugin, please star it on GitHub — it shows me that the plugin is useful to you and motivates me to keep developing it.

🐛 If you find a bug or would like to request a feature, open a GitHub issue in any language — I will review your proposal and implement useful suggestions in a future plugin version.

⚡ Overview

dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.

graph LR
    subgraph Client [Browser Web UI]
        Mic[🎙️ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
        Wave[🌊 Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
    end

    subgraph Host [DSH Host Backend]
        Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
        PTT --> FFMPEG
        FFMPEG --> Chain{Fallback Chain}
        
        Chain -->|1st Priority| P1[Deepgram / Nova-2]
        Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
        Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
    end

    subgraph Output [Target]
        P1 --> Composer[💬 Web Composer / Chat]
        P2 --> Composer
        P3 --> Composer
    end

    style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
    style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4

✨ Key Features

  • 🎙️ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (vadSilenceMs, default 700ms) and typed into the composer in real time.
  • 🌊 Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (autoSendMs, default 4000ms).
  • 🎮 Tactile Push-to-Talk:
    • Mouse: Hold the wave button — releasing sends the message; dragging pointer away discards.
    • Keyboard: Hold Ctrl (or custom hotkey) for hands-free speaking; press Esc to cancel.
  • ⚡ Zero-Latency In-Browser Captions (browser): Chrome Web Speech API recognition runs 100% locally with live floating captions as you speak.
  • 🛡️ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
  • 🧠 Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
  • 🎵 Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
  • 🔇 Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
  • 📊 Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
  • 🔒 Zero API Key Leakage: Keys are resolved on the host via ctx.credentials (credentialRef) and never transmitted to browser clients.
  • 🖥️ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (whisper-server) with on-the-fly ffmpeg transcode.
  • ⚡ SenseVoice-ONNX / Sherpa-ONNX (0.8.11, updated 0.8.32): Ultra-fast (~50–100ms) non-autoregressive local STT engine with 1-click automatic model installation to ~/.dsh/models/sensevoice.
  • 🎙️ Gated Turn-Taking & Echo Prevention (0.8.32): Automatically mutes the mic while the assistant is speaking (dsh:tts:start/stop) and supports speech-triggered Barge-In.
  • 👻 Live Ghost Interim Preview (0.8.32): Visual real-time preview of spoken words before final chunk transcription.
  • ⏱️ SVG Silence Ring Timer (0.8.32): Circular animated countdown indicator in the composer dock during the pending message delay.
  • 🗣️ Hands-free Voice Actions (0.8.32): Spoken commands ("send", "cancel", "clear", "new line") trigger UI actions directly.
  • 🧩 Structured Prompt Voice Selection (0.8.32): Spoken options automatically select buttons in interactive assistant choice prompts.
  • 📖 Developer Lexicon & IT Jargon Correction (0.8.32): Phonetic normalization for developer slang (GitHub, Docker, Kubernetes, pnpm) with customizable settings dictionary.
  • 🌐 Realtime Audio Streaming (0.8.11): Low-latency WebSocket bridge (/dsh-voice/realtime) for OpenAI Realtime API or local Sherpa-ONNX streaming. API keys stay securely on the host.
  • 🌊 Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.

📝 What's New in 0.8.32

10-point feature epic & SenseVoice 1-Click Installer (Issue #129):

AreaFeatureDescription
Local ASR1-Click SenseVoice-SmallAutomated download and extraction of model.int8.onnx and tokens.txt directly to ~/.dsh/models/sensevoice via host loopback endpoint.
Turn-TakingGated ModeMicrophone input is muted when assistant audio starts (dsh:tts:start) and unmuted on dsh:tts:stop, eliminating acoustic feedback.
InterruptionBarge-InSpeaking immediately signals assistant TTS to pause and abort current speech playback.
FeedbackGhost Interim PreviewSemi-transparent live text preview in composer pill during recognition before sentence finalizing.
CountdownSVG Silence RingCircular SVG progress ring visually ticking down pending auto-send delay.
ControlHands-free ActionsSpoken control words ("send", "cancel", "clear", "new line") trigger actions instead of becoming message text.
PromptsStructured PromptsSpoken replies automatically match and submit options in active harness interactive prompts.
VocabularyIT Jargon NormalizerAuto-corrects spoken developer slang to canonical spelling (GitHub, Docker, Kubernetes, etc.) with custom dictionary in Settings.
ConfigurationIndividual TogglesDedicated switches for every enhancement in Settings card with native --dsw-alias-* token styling.

📝 What's New in 0.8.19

Quality batch after the v0.8.18 review (Gitea #79–#87, PR #88):

AreaChange
Settings placeholdersModel/path hints resolve through locale at render time — no frozen i18n keys (#79)
Model field hintswhisperModel / sensevoiceModel show path-to-model copy, not binary-autostart hints (#82)
VisualizerLiquid wave / dynamic orb read DSH theme tokens instead of hardcoded hex (#80)
Style isolationInjected CSS uses data-dsh-plugin="dsh-voice" so neighbour HMR cleanup cannot strip styles (#81)
Source languageCode comments, errors, and tests are English; EN edit-command phrases added. Spoken RU STT patterns remain for recognition (#86)
LocalesChanged: the full inline ru dictionary was removed. English is the only bundled locale. Install the translation plugin for Russian UI (#85)
Client sourceBrowser client is built from ordered lib/client-src/*.js fragments via npm run build:client (#87)
Process docsAdded docs/design/DESIGN.md, index.md, project AGENTS.md, docs/testing/unit.md (#83)

[!IMPORTANT] v0.8.19 locale behavior: without a translation plugin the Web UI stays in English. Session/edit spoken command phrases still match Russian speech for STT where documented; interface labels do not ship a second dictionary.


🎮 Four Ways to Speak

ModeGesture / TriggerBehavior
DictationClick 🎙️ MicSpeech is sliced on pauses (vadSilenceMs) and typed live into composer
Voice MessageClick 🌊 WaveRecords until stopped, then sends after cancel window (autoSendMs)
Mouse PTTHold 🌊 WaveRecords while held; release sends message, drag off button to discard
Keyboard PTTHold CtrlHands-free recording; release sends message, press Esc to discard

[!TIP] You can customize the keyboard modifier in settings (hotkey: Control, Alt, Shift, or any KeyboardEvent.code).


🛠️ Supported Providers Matrix

Provider KeyService BackendDefault ModelCredential RefFeatures & Notes
browserWeb Speech APINative BrowserNoneZero latency, floating live captions in Chrome
deepgramDeepgram APInova-2DEEPGRAM_API_KEYUltra-fast cloud transcription
groqGroq Whisperwhisper-large-v3-turboGROQ_API_KEYNear-instant inference speed
hfHuggingFace Inferenceopenai/whisper-large-v3HF_TOKENHigh-accuracy open Whisper
local-whisperLocal whisper.cppServer definedNone100% private, offline, no internet needed
sensevoiceSenseVoice-ONNX / Sherpa-ONNXSenseVoiceSmallNoneUltra-fast (~50ms) local non-autoregressive STT

🚀 Ready-Made Presets (Plug & Play)

Just specify the name in your fallback chain and add the corresponding API key:

  • openai (whisper-1) → OPENAI_API_KEY
  • siliconflow (SenseVoiceSmall) → SILICONFLOW_API_KEY
  • mistral (voxtral-mini-latest) → MISTRAL_API_KEY
  • openrouter (google/gemini-2.5-flash) → OPENROUTER_API_KEY
  • deepinfra (whisper-large-v3-turbo) → DEEPINFRA_API_KEY
  • fireworks (whisper-v3-turbo) → FIREWORKS_API_KEY

📦 Quick Installation

dsh plugin --profile web add @goodandready/dsh-voice

[!IMPORTANT] Restart DSH Web UI after installation (systemctl --user restart dsh-web) and refresh your browser tab.


⚙️ Configuration

Open Settings → Plugins → Plugin settings → Voice in the Web UI:

- id: dsh-voice
  config:
    dictation:
      language: ru
      vadSilenceMs: 700
      chain:
        - provider: deepgram
        - provider: groq
        - provider: local-whisper
    message:
      language: ru
      autoSendMs: 4000
      chain:
        - provider: openai
        - provider: local-whisper
    hotkey: Control
    autoStart: true
    whisperModel: /models/ggml-medium-q8_0.bin

🤖 Agent Tool & HTTP API

Agent Tool (transcribe_audio)

Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk.

Internal HTTP Endpoints

  • POST /dsh-voice/transcribe — { dataBase64, mimeType, mode } → { ok, text, provider, tookMs }
  • POST /dsh-voice/polish — { text } → { ok, text }
  • GET /dsh-voice/status — Returns daemon status, active chains, SenseVoice and realtime config.
  • GET /dsh-voice/realtime — WebSocket upgrade for low-latency audio streaming (OpenAI Realtime API / Sherpa-ONNX). Accepts binary audio chunks, returns JSON text deltas.

📄 License

MIT © GooDAnDReaDY