Instant interrupt
Audio is streamed in ~50 ms slices, so pressing Esc cuts the assistant off mid-word and returns to listening in under 100 ms.
Available for new projects
Real-time voice, screen vision, desktop automation and persistent memory — packaged into one JARVIS-style assistant that runs on your machine. Below is the full case study of the system I built.
Case study
A real-time voice AI that hears, sees, understands and controls a computer. It streams native audio over the Gemini Live API, sees the screen and webcam, remembers your projects and preferences between sessions, and operates apps the way a person would — with hands, not just APIs.
The brief was to remove every point of friction: silence while it thinks, double responses, slow interrupts, search results that link to homepages. The result feels like talking to a colleague, not a chatbot.
Audio is streamed in ~50 ms slices, so pressing Esc cuts the assistant off mid-word and returns to listening in under 100 ms.
An LLM-grounded search and a news search race in parallel threads. First valid answer wins — real article links, never homepages.
Exponential back-off reconnects (3 → 6 → 12 → 60 s) with clear status, and full state reset on every new session.
Capabilities
Ultra-low-latency conversation in any language, with automatic language detection and correct forms of address.
Launch apps, manage windows, volume, brightness, Wi-Fi, shortcuts and power — by voice.
Screen capture and webcam piped into the live session. “What’s on my screen?” just works.
Remembers projects, preferences and personal context across sessions and adapts its briefings.
News, research, price, compare — grounded answers with sources, with DuckDuckGo as a safety net.
Give it a multi-step goal — it plans, acts, checks the result and corrects itself.
Compose and send messages, talk to the assistant remotely, and get alerts on your phone.
Live CPU, RAM, GPU and temperature monitoring with spoken alerts when thresholds are crossed.
OS-native scheduled notifications and named command routines you can run or schedule daily.
Live demo
A simulated session showing how a request flows from voice to tool call to result. Pick a scenario.
Architecture
When one method can’t reach a control, the assistant falls back to the next. Every action is logged with app, window, target, method, result and timing.
Screenshot + multimodal model finds the element visually
Synthesised input when UI Automation can’t reach a control
Find, click, type, read and select controls by name, id or class
Process and window enumeration, focus, move, resize, close
Downloads
Grab the Windows installer. You’ll need your own API keys — see the FAQ.
Windows may show a SmartScreen warning on first launch because the installer isn't code-signed yet — this is standard for new, independently-built apps, not a sign the file is unsafe. Click "More info" → "Run anyway" to continue. Download only from this page.
Windows 10/11 · Download only from this page and check the file size before running.
Process
You describe the workflow and the tools it must control. I reply with scope, timeline and a plan.
Core loop first, then tools one by one — each verified against real apps, not mocks.
Privacy gates, confirmations for risky actions, reconnects, logging and auto-start.
Installer, documentation and a walkthrough call so it’s yours to run and extend.
FAQ
The assistant runs on your computer. Voice reasoning uses a cloud model API by default, and an offline fast path plus an optional local LLM (Ollama) cover basic commands without internet.
Windows, macOS and Linux for the core assistant. The deepest desktop-control layer (Win32 + UI Automation) is Windows-first.
Capabilities such as camera, screen, files, browser and messaging have individual privacy switches that are enforced before a tool runs, and destructive actions can require explicit confirmation.
Yes — new tools are registered with a schema and routed through the same permission layer, so custom integrations are a natural extension.
Your use-case, the apps you want it to control, and your own API keys. Keys stay in a local config on your machine.
The installer isn't code-signed yet, so Windows SmartScreen shows a caution screen the first time — this is standard for new, independently-built apps and not a sign of a virus. Choose "More info" → "Run anyway". Only download it from this page.
Tell me what you want it to do. I’ll reply within 24 hours with a plan.