Available for new projects

I build AI assistants that hear.

Real-time voice, screen vision, desktop automation and persistent memory — packaged into one JARVIS-style assistant that runs on your machine. Below is the full case study of the system I built.

  • 0voice-callable tools
  • 0interrupt latency
  • 0operating systems
  • 0control fallback layers
PythonGemini Live APIClaudePyQt6PlaywrightFastAPIRustTypeScriptTelegramLiveKitOllamaWhisperWin32 / UIA PythonGemini Live APIClaudePyQt6PlaywrightFastAPIRustTypeScriptTelegramLiveKitOllamaWhisperWin32 / UIA

Case study

JARVIS AI
a desktop assistant that behaves like software from a movie.

A real-time voice AI that hears, sees, understands and controls a computer. It streams native audio over the Gemini Live API, sees the screen and webcam, remembers your projects and preferences between sessions, and operates apps the way a person would — with hands, not just APIs.

The brief was to remove every point of friction: silence while it thinks, double responses, slow interrupts, search results that link to homepages. The result feels like talking to a colleague, not a chatbot.

  • Hybrid voice + keyboard input with a live HUD (waveform, log, camera feed)
  • Local Fast Path — simple commands still work when the cloud is down
  • Offline chat fallback through a local Ollama model
  • Per-capability privacy switches enforced before a tool runs

Instant interrupt

Audio is streamed in ~50 ms slices, so pressing Esc cuts the assistant off mid-word and returns to listening in under 100 ms.

Parallel news search

An LLM-grounded search and a news search race in parallel threads. First valid answer wins — real article links, never homepages.

Self-healing sessions

Exponential back-off reconnects (3 → 6 → 12 → 60 s) with clear status, and full state reset on every new session.

Capabilities

Everything a personal AI should do,
in one voice-driven core.

Real-time voice

Ultra-low-latency conversation in any language, with automatic language detection and correct forms of address.

System control

Launch apps, manage windows, volume, brightness, Wi-Fi, shortcuts and power — by voice.

Visual awareness

Screen capture and webcam piped into the live session. “What’s on my screen?” just works.

Persistent memory

Remembers projects, preferences and personal context across sessions and adapts its briefings.

Multi-mode web search

News, research, price, compare — grounded answers with sources, with DuckDuckGo as a safety net.

Autonomous tasks

Give it a multi-step goal — it plans, acts, checks the result and corrects itself.

Telegram & messaging

Compose and send messages, talk to the assistant remotely, and get alerts on your phone.

Hardware telemetry

Live CPU, RAM, GPU and temperature monitoring with spoken alerts when thresholds are crossed.

Reminders & routines

OS-native scheduled notifications and named command routines you can run or schedule daily.

Live demo

Talk to it.
Watch it work.

A simulated session showing how a request flows from voice to tool call to result. Pick a scenario.

jarvis — simulated session IDLE

Architecture

A four-layer control ladder.
It never gives up on the first try.

When one method can’t reach a control, the assistant falls back to the next. Every action is logged with app, window, target, method, result and timing.

04

Vision

Screenshot + multimodal model finds the element visually

last resort
03

Keyboard & mouse

Synthesised input when UI Automation can’t reach a control

fallback
02

UI Automation

Find, click, type, read and select controls by name, id or class

precise
01

Native Win32

Process and window enumeration, focus, move, resize, close

fastest
Voice / text
LLM reasoning
Tool router + privacy gate
Action
Spoken result

Process

From idea to a working assistant.

  1. 01

    Brief

    You describe the workflow and the tools it must control. I reply with scope, timeline and a plan.

  2. 02

    Build

    Core loop first, then tools one by one — each verified against real apps, not mocks.

  3. 03

    Harden

    Privacy gates, confirmations for risky actions, reconnects, logging and auto-start.

  4. 04

    Deliver

    Installer, documentation and a walkthrough call so it’s yours to run and extend.

FAQ

Questions, answered.

Does it run locally or in the cloud?

The assistant runs on your computer. Voice reasoning uses a cloud model API by default, and an offline fast path plus an optional local LLM (Ollama) cover basic commands without internet.

Which operating systems are supported?

Windows, macOS and Linux for the core assistant. The deepest desktop-control layer (Win32 + UI Automation) is Windows-first.

Is it safe to let an AI control my PC?

Capabilities such as camera, screen, files, browser and messaging have individual privacy switches that are enforced before a tool runs, and destructive actions can require explicit confirmation.

Can you connect it to my own tools or APIs?

Yes — new tools are registered with a schema and routed through the same permission layer, so custom integrations are a natural extension.

What do I need to provide?

Your use-case, the apps you want it to control, and your own API keys. Keys stay in a local config on your machine.

Why does Windows warn me when I install it?

The installer isn't code-signed yet, so Windows SmartScreen shows a caution screen the first time — this is standard for new, independently-built apps and not a sign of a virus. Choose "More info" → "Run anyway". Only download it from this page.

Ready to build your own JARVIS?

Tell me what you want it to do. I’ll reply within 24 hours with a plan.

…or leave a request right here