wilsonix-studio
AI-powered desktop DAW - stem separation, chord detection, karaoke
- Stars
- 0
- Language
- —
- Created
- Sep 1, 2026
- Updated
- Sep 4, 2026
Introduction
WilSonix Studio PRO
AI-Powered Desktop DAW — Stem Separation, Karaoke & Live Performance Suite
Separate. Align. Master. Perform.
Download
Windows Installer — Download v0.2.4 from Releases
| Platform | Status |
|---|---|
| Windows x64 (NSIS Installer) | Stable (v0.2.4) |
| Android (Capacitor APK) | Beta |
| macOS / Linux | Build from source |
What's New in v0.2.4
- Stem-Targeted Harmonies: Generate 3-part diatonic harmonies on either
LEAD_VOCALSorBACKGROUND_VOCALSdirectly from the mixer channel strip. - Dynamic Stem Indicator: Vocal Studio HUD displays the exact active target stem and musical key.
- Go Gateway Proxy Fix: Direct reverse-proxy forwarding through port 7860 for all harmony, spectral repair, and vocal character endpoints.
- Continuous Phase Vocoder Auto-Tune: Zero audio slicing, zero chunk boundaries, zero boundary clicks, and zero comb-filtering.
- High-Speed 3-Part Vocal Harmonizer: Generates +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems in seconds.
- Adaptive Light/Dark Mode: Vocal Studio and Spectral Repair modals dynamically restyle for both light and dark themes.
- Consonant & Sibilance De-Ess Guard: Prevents sizzling/sparkling phase fuzz by automatically detecting and bypassing pitch-shifting on unvoiced consonants (s, sh, t, k, p, breaths), keeping them 100% pure acoustic audio.
- AI Singer Gender & Duet Detection: Real-time vocal acoustic classifier detects Male, Female, and Duet/Polyphonic vocal intervals.
- 3-Tier Color-Coded Karaoke Teleprompter: Words dynamically light up in Electric Neon Cyan (Male), Radiant Hot Pink (Female), and Golden Amber (Duet / Both), accompanied by active singer badges on the TV stage prompter (
[♂ SINGER 1],[♀ SINGER 2],[👥 DUET / BOTH]). - Interactive Singer Override: Clickable
[♂ HE] / [♀ SHE] / [👥 DUET]pills on every lyric line in the teleprompter editor allow instant manual reassignment with automatic saving. - 1080p MP4 Video Produce with Burned Multi-Style Subtitles: Advanced SubStation Alpha (
.ass) rendering engine burns exact singer colors directly into exported videos via FFmpeg. - Continuous Pitch Performance Scoring: Overhauled scoring engine evaluates pitch accuracy and rhythm on 50ms hop windows across the full vocal take with cents-level vibrato tolerance (no hardcoded scores).
- AI Spectral Repair & Stem "Healing Brush": Interactive 2D STFT spectrogram viewer ($20\text{ Hz} - 20\text{ kHz}$) with click-and-drag marquee box selection and 3 surgical inpainting algorithms: Attenuate (-18dB), Ambient Fill (Wiener), and Harmonic Interpolation.
- Next-Gen Vocal Studio & Diatonic 3-Part Harmonizer: Generates discrete +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems locked to the detected song key and chords, adding them directly into the DAW mixer as controllable faders.
- Formant-Locked Voice Character Morphing: Apply Female Bright, Male Warmth, or Vintage Radio profiles without chipmunk distortion.
- Native Rust DSP Core (
crates/wilsonix-dsp): SIMD RealFFT inpainting engine and GCC-PHAT transient phase alignment for zero-latency C-speed audio computation.
Features
Stem Separation
Upload any audio or video file and let AI split it into isolated stems:
| Engine | Stems | Speed | Best For |
|---|---|---|---|
| Demucs v4 | 6 (Vocals, Drums, Bass, Guitar, Piano, Other) | ~2.5 min | Full studio decomposition |
| Demucs v4 | 4 (Vocals, Drums, Bass, Other) | ~1.5 min | Classic multi-track |
| UVR MDX-Net ONNX | 2 (Vocal + Backing) | ~50 sec | Fast karaoke / minus-one |
| BS-RoFormer | 2 (Vocal + Backing) | SOTA (12.98 dB SDR) | Studio mastering quality |
- 32-bit float processing throughout
- Multi-guitar spatial decomposition (4-way azimuth)
- Solo Protect — recovers lead instruments from vocal stem
- Quality modes: Draft / Studio HQ / Ultra HD / Mastering Grade
DAW Mixer
Full-featured multi-channel mixer with per-stem controls:
- Volume Fader (0–150%) + Pan Knob (stereo panner)
- Mute / Solo per channel
- 3-Band Parametric EQ — Low shelf (250 Hz), Mid peaking (1.2 kHz), High shelf (4.5 kHz)
- Low-Cut Filter — 80 Hz rumble removal
- Compressor — Threshold -18 dB, ratio 3:1
- Tape Saturator — Analog tube warmth, 4x oversampling
- Stereo Chorus — LFO-modulated delay
- Haas Effect Widener — Stereo spatial expansion
- Reverb Send — Convolution hall bus (2.2s)
- Delay Send — 350ms tempo-synced feedback delay
- Per-channel FFT analyser + OLED waveform display
- 6 built-in FX presets per channel
- Import custom WAV files as mixer channels
Chord Detection
- Real-time chord recognition powered by AI
- Instrument selector: Guitar / Bass / Piano / All (Mix)
- Scrollable chord progression timeline in the Harmonics strip
- Chords cached in database for instant reload
Karaoke Stage
- Whisper AI word-level transcription (millisecond precision)
- Multi-language: English, Bisaya, Tagalog, Japanese, Korean, Chinese, Spanish, French, German
- OLED teleprompter with physics bouncing ball + word-by-word glow highlight
- 6 trance visual themes (Cyber Aurora, Warp Tunnel, Plasma Orbs, Synthwave Grid, Matrix Starfield, Kaleidoscope Nebula)
- Lyrics editor with undo/redo, gap auto-fill, hallucination detection panel
- Export: .LRC and .ASS subtitle files
- Sync offset nudge (-0.1s / +0.1s)
- Live microphone sing-along with real-time pitch tracking
- Pitch radar canvas visualization
Auto-Tune & Vocal Processing
- Neural Auto-Tune (WORLD Vocoder) — formant-preserving pitch correction
- Scales: Chromatic, Auto, Major keys
- Presets: T-Pain, Pop, Natural
- AI Mic Enhancement — Denoise + De-Ess + Capsule Exciter
- Profiles: Neumann U87, Shure
Dual-Screen TV Stage
- Pop-out dedicated fullscreen teleprompter to any external TV or monitor
- Zero-latency lyrics teleprompter with chord sync
- Perfect for live karaoke performance
Karaoke Video Production
- 1080p MP4 export with burned-in animated glowing lyrics
- Highlight colors: Neon Cyan, Warm Amber, Vibrant Pink, Emerald Mint
- Vocal guide volume slider
- Title / Artist metadata
- Performance video remux (browser recording → broadcast H.264/AAC)
Auto-Mastering
- One-click AI mastering to -14 LUFS (streaming standard)
- ITU-R BS.1770 / EBU R128 K-weighting filter
- Multi-band DSP processing
- True-Peak limiter
- 24-bit mastered WAV output
File Inspector
- MIDI file parser — notes, tempo map, piano roll visualization
- ASS subtitle inspector — styles, events, karaoke timings
- Quick ASS override tag parser
Backend Logs
- Real-time log streaming with color-coded output
- Filter by level: All / Errors / Warnings / Info
- Auto-scroll with clear button
Library & Project History
- Search and filter past projects
- Status tracking (completed / failed / processing)
- One-click "Open in Studio" to reload stems
- Download stems as ZIP
- Auto-refresh polling for active jobs
Admin Control Center
- CPU, RAM, GPU, Disk telemetry
- User management (roles: admin / producer)
- Job queue monitoring
- Stem purge for expired data
Tech Stack
| Layer | Technology | Purpose |
|---|---|---|
| Desktop Shell | Tauri v2 (Rust) | Native window, system integration, process management |
| Backend Server | Go | HTTP server, WebSocket relay, SQLite DB, auth, routing |
| AI Engine | Python 3.12 + PyTorch | Stem separation, chord detection, lyrics transcription |
| ML Models | Demucs v4, UVR MDX-Net, BS-RoFormer, OpenAI Whisper | AI audio processing |
| Frontend | Vanilla JS + Tailwind CSS + Anime.js + Lucide Icons | Reactive DAW workspace |
| Audio DSP | Web Audio API (32-bit float) | Real-time mixing, EQ, compression, effects |
| Database | SQLite (via Go) | Users, jobs, stems, lyrics, chords |
| Mobile | Capacitor + ONNX Runtime Mobile | Android APK with offline inference |
| Installer | NSIS | Windows installer packaging |
| Build | PyInstaller | Python → single .exe bundling |
System Requirements
| Component | Minimum | Recommended |
|---|---|---|
| OS | Windows 10 x64 | Windows 11 |
| RAM | 4 GB | 8 GB+ |
| Storage | 2 GB free | 5 GB+ (for models) |
| CPU | Dual-core 2 GHz | Quad-core+ with AVX |
| GPU | None (CPU works) | NVIDIA GPU with CUDA for faster inference |
Supported Formats
| Input | Formats |
|---|---|
| Audio | WAV, FLAC, MP3, M4A, OGG, OPUS |
| Video | MP4, MKV, MOV, WebM, AVI |
| Subtitles | .ASS, .SSA |
| MIDI | .mid, .midi |
Keyboard Shortcuts
| Key | Action |
|---|---|
Space | Play / Pause |
M | Mute selected channel |
S | Solo selected channel |
0 | Rewind to beginning |
[ / ] | Set Loop Point A / B |
L | Toggle Loop |
Ctrl+Z | Undo |
Ctrl+Shift+Z | Redo |
Esc | Close modal |
Mobile (Android)
WilSonix Studio PRO is available as an Android APK via Capacitor:
- Swipeable 6-stem mixer carousel
- Gesture-locked audio sliders
- Offline AI inference with ONNX Runtime Mobile
- Touch-optimized UI with safe-area support
Building from Source
Prerequisites
- Go 1.22+
- Rust + Tauri CLI
- Python 3.12 with venv
- NSIS (for Windows installer)
Build Steps
# Clone
git clone https://github.com/ewceniza9009/pygo.git
cd pygo
# Python environment
python -m venv python_env
python_env\Scripts\activate
pip install -r requirements.txt
# Build Python worker (PyInstaller)
pyinstaller --onefile --name python_worker --noconfirm \
--distpath src-tauri\resources \
--collect-data audio_separator \
--collect-submodules audio_separator \
--collect-all onnxruntime \
run_worker.py
# Build Go server
go build -ldflags "-s -w" -o src-tauri\resources\pygo_server.exe .
# Build Tauri desktop app
cd src-tauri
cargo tauri build
Architecture
┌─────────────────────────────────────────────┐
│ Tauri (Rust) Shell │
│ ┌──────────┐ ┌──────────────────────────┐ │
│ │ Webview │ │ Process Manager │ │
│ │ (HTML) │ │ ├─ pygo_server.exe │ │
│ │ │ │ └─ python_worker.exe │ │
│ └──────────┘ └──────────────────────────┘ │
└─────────────────────────────────────────────┘
│ HTTP/WebSocket │ HTTP
┌────▼────┐ ┌────▼────┐
│ Go │────────────│ Python │
│ Server │ proxy WS │ FastAPI │
│ :7860 │ │ :19876 │
└────┬────┘ └────┬────┘
│ │
┌────▼────┐ ┌────▼────────────┐
│ SQLite │ │ PyTorch / ONNX │
│ (DB) │ │ Demucs/Roformer │
└─────────┘ │ Whisper │
└─────────────────┘
Developer
Erwin Wilson E. Ceniza
License
Proprietary. All rights reserved.