Skip to content

About

Computa, control this computer, thank you.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

voice-control

A background macOS agent that listens for "computa, …" and turns the command that follows into an HTTP request, an OBS change, a media key, a Spotify Web API control, an audio device switch, a Home Assistant service call, or a local program.

computa, mute            ->  POST :8009/v1/voice/canary/mute/on
computa, unmute          ->  POST :8009/v1/voice/canary/mute/off
computa, deafen          ->  POST :8009/v1/voice/canary/deaf/on
computa, push to talk    ->  POST :8009/v1/voice/canary/mode/PUSH_TO_TALK
computa, desk mic        ->  POST :8080/v1/input/set/desk
computa, camera          ->  OBS scene "Camera Screen"
computa, main            ->  OBS scene "Main Screen"
computa, ps5             ->  OBS scene "Main Screen", then animate the
                             "Cam Link Screen" source in (or out)
computa, hide ps5        ->  animate "Cam Link Screen" out
computa, skip            ->  media key: next track
computa, play            ->  media key: play/pause
computa, music quieter   ->  Spotify: lower active device volume 10%
computa, headphones      ->  default output device: AirPods
computa, speakers        ->  default output device: the desk speakers
computa, coffee          ->  Home Assistant: switch.coffee_machine
                             toggle
computa, are you         ->  a tone back, and nothing else
  listening?

Targets are discord-rpc-control on localhost and mic-api on the Linux box. Nothing is hardcoded — commands are a TOML table, so adding one is a config edit and a restart.

Discord RPC mutes inside Discord and leaves the OS input device open, which is what makes "computa, unmute" work while muted. Do not swap the dispatch for an OS-level mute; the daemon would go deaf to its own un-mute.

How it works

cpal capture ──► resample 48k→16k mono ──► 3s pre-roll ring
                          │
                    ┌─────┴─────┐
                    │  IDLE     │  openWakeWord scores every 80ms hop
                    └─────┬─────┘
                    "computa" ✔ → chirp
                          │
                    ┌─────┴─────┐
                    │ LISTENING │  Silero VAD: up to 2.5s for the
                    └─────┬─────┘  command to start, then 700ms of
                          │        silence ends it
                          │        (seeded with 900ms of pre-roll)
                          │
                  whisper.cpp base.en  →  "computa mute"
                          │
                  normalise → suffix candidates → fuzzy match
                          │
                  commands.toml → POST … → ok/fail tone

Two stages rather than one because a wake-word model is cheap enough to run continuously and whisper is not.

The pre-roll is the part that is easy to get wrong. "computa, mute" is one breath, so the command is already partly spoken when the wake word starts scoring — and the detector then needs patience more hops before it commits. Between them the whole utterance can be in the past by the time the detector says anything, so the window reaches back far enough to recover the wake word too. That the wake word ends up in the clip is fine: the matcher works over suffixes and drops it.

It also means the capture starts out having already heard speech — the wake word — whatever the user does next. Ending it on trailing silence alone therefore closes the window about 700 ms after "computa", which is not long enough to say "computa", think for a beat, and then give the command: the first syllable arrives after the capture has already been sent to whisper. So silence does not end anything until something has been said live, past the end of the pre-roll. Until then the only clock running is grace_ms, and when that runs out the wake word is written off as a false trigger and dropped quietly.

Nothing is lost in the brisk case. Said in one breath, "computa mute" is still being spoken when the detector commits, so its tail arrives live, and the hangover ends the capture as promptly as it ever did — the grace period only applies to a capture where nothing has been said yet, and there is nothing to cut short.

[listen]
grace_ms = 2500     # how long the command may still be starting
silence_ms = 700    # trailing silence that ends one already started
max_ms = 8000       # ceiling on a capture, pre-roll included

Both stages sit behind traits (wake::Detector, stt::Transcriber), and the wake word is openWakeWord through that trait, so the engine is one file: src/wake/oww.rs.

It is three ONNX graphs in series, scored once per 80 ms hop:

1280 samples ─► melspectrogram ─► 8 mel frames of 32 bins
                                     │  (ring of 80)
                       last 76 frames ▼
                                  embedding ─► 96 floats
                                     │  (ring of 16)
                                     ▼
                                   wake ─► one probability

The first two are openWakeWord's own pretrained feature extractors and are the same whatever the wake word is; only the third is trained per word. That split is the whole reason this replaced rustpotter, which matched a waveform against a dozen recordings of it and could not tell the wake word from a television. The embedding model is a speech representation trained on a very large corpus, so the classifier is asked "what was said", not "how close is this to my takes" — a working model reads ~0.00 on speech that is not its word rather than 0.6.

Inference runs on ONNX Runtime via ort, which was already in the tree: the Silero VAD uses it too.

Setup

Needs cmake (brew install cmake) to build whisper.cpp. The first build takes a few minutes; that is whisper.cpp compiling, not a hang.

1. Whisper model

mkdir -p ~/.config/voice-control
curl -L -o ~/.config/voice-control/ggml-base.en-q5_1.bin \
  https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en-q5_1.bin

base.en is noticeably better than tiny.en on one- and two-word commands and still runs in well under 200 ms on Apple Silicon. Drop to ggml-tiny.en-q5_1.bin if you want the latency back.

2. Wake word model

Three files: the two shared feature extractors, which never change, and a classifier for the word itself.

cd ~/.config/voice-control
base=https://github.com/dscripka/openWakeWord/releases/download/v0.5.1
curl -L -O $base/melspectrogram.onnx
curl -L -O $base/embedding_model.onnx
curl -L -O $base/alexa_v0.1.onnx && mv alexa_v0.1.onnx alexa.onnx

alexa, hey_jarvis, hey_mycroft and hey_rhasspy are the pretrained words, and they are very good: alexa_v0.1 reads 1.000 on "alexa" and below 0.07 on everything else tried here, including real speech, "computer", and a directory of takes of a different wake word. Point [wake] model at whichever one you want.

Any other word has to be trained, which openWakeWord does with a corpus of synthetic speech, not a handful of takes — see its training notebooks. A model trained on too few negatives is worse than useless here: one tried during this port fired at 0.98+ on ordinary dialogue that sounded nothing like its word, which is exactly the failure that made rustpotter unusable. Check any model you are given before trusting it:

# Should all fire.
cargo run --release -- score positives/*.wav
# Should all read ~0.00. If they do not, the model is not usable,
# whatever it scores on the word itself.
cargo run --release -- score negatives/*.wav

3. Commands

cp commands.example.toml ~/.config/voice-control/commands.toml

Set mic under [targets] to the Linux host, and check the Discord branch in the discord URL matches DISCORD_BRANCH in discord-rpc-control (its installer defaults to canary).

4. Tune the threshold

cargo run --release -- listen

Prints an input level once a second and a line per detection:

  level 0.067  ###
  1  alexa  score 0.998

If the level stays at zero, the microphone permission is the problem, not the model — it says so after five seconds of silence.

With a well-trained model there is not much to tune: scores sit at either end of the range, so anything from 0.3 to 0.8 behaves the same and 0.5 is fine. Tuning only starts to matter when the model is marginal, and then the honest answer is usually a better model.

Guessing at it from a live microphone is slow, so there are two subcommands for doing it from recordings instead. record writes the microphone to a directory of one-minute wavs and prints every hit as it happens, with the file and offset:

cargo run --release -- record ~/negatives     # leave it running
  1  audio-0003.wav    12.4s  score 0.981     # ← go and listen to that

Leave it going through whatever sets the daemon off. Then score replays wavs through the detector — no microphone, no whisper, nothing dispatched — and prints the peak score, the longest run of hops over the threshold, and whether it would have fired:

cargo run --release -- score ~/negatives/*.wav
audio-0003.wav  peak 0.981 at 12.4s  run 6  FIRED at 12.2s

threshold wants to sit in the gap between what your voice scores and what the room scores. If there is no gap — if run is as long on television as it is on you — no threshold will fix it.

patience is the other lever: how many consecutive 80 ms hops have to clear the threshold before it counts. It costs patience × 80 ms of latency, which the pre-roll absorbs, and it drops false positives that only glance off the model for a frame or two. Two is a good default; raising it trades a little responsiveness for fewer of them.

5. Install

./scripts/install-agent.sh

Builds release, installs to ~/Library/Application Support/voice-control, ad-hoc signs the binary (so the microphone grant survives rebuilds), and bootstraps the LaunchAgent.

macOS will ask for microphone access the first time it runs. If no prompt appears and it never hears anything, check System Settings → Privacy & Security → Microphone.

Media key commands need a second grant, Accessibility, which is never prompted for — add ~/Library/Application Support/voice-control/voice-control under System Settings → Privacy & Security → Accessibility by hand. Skip it if you have no media commands. Both grants are keyed to the code signature, which is why the script ad-hoc signs the binary before launchd sees it: without that, every rebuild would look like a brand new app and drop them.

Usage

voice-control                    run the daemon
voice-control devices            list audio devices and their aliases
voice-control output headphones  switch the default output, bare
voice-control input "Wireless"   switch the default input, bare
voice-control listen             wake word detections only
voice-control transcribe f.wav   STT + matcher against a file
voice-control run "computa ps5"  match a phrase and dispatch it
voice-control replay f.wav       whole pipeline against a file
voice-control hass               list home assistant entities
voice-control hass speaker       the ones matching a filter
voice-control hass switch.desk_speakers      call a service, bare
voice-control obs                list obs scenes
voice-control obs "Camera Screen"  switch to a scene
voice-control obs sources        list the sources in the current scene
voice-control obs filters "Main Screen"      list a scene's filters
voice-control obs toggle "Cam Link Screen"   flip a source, bare
voice-control media next         press a media key, bare
voice-control spotify authorize  authorize Spotify once with PKCE
voice-control spotify devices    list Spotify Connect device ids
voice-control spotify volume 40  set Spotify's volume, bare

transcribe is the fast way to grow phrase lists without talking to a microphone. run takes the phrase directly and dispatches it exactly as the daemon would, sequences and waits included — the way to check what a command does once you know it matches. replay runs a file through wake word, VAD, STT, matching and dispatch — everything except opening the microphone — so you can check a change end to end. Generate test clips with say -v Samantha -o x.aiff "computa mute" piped through afconvert -f WAVE -d LEI16@16000 -c 1 x.aiff x.wav.

OBS scenes

Scene switching goes over obs-websocket 5.x. Enable the server in OBS under Tools -> WebSocket Server Settings, then:

[obs]
url = "ws://127.0.0.1:4455"   # wss:// for an OBS on another machine
password = ""                 # empty reads OBS_PASSWORD instead

[[commands]]
name = "camera scene"
phrases = ["camera", "cam", "show me"]
scene = "Camera Screen"

Scenes are enumerated one command at a time on purpose: the daemon can only switch to scenes you named, never to whatever it thought it heard. scene is matched by OBS exactly, so copy the names from:

voice-control obs                    # list them
voice-control obs "Camera Screen"    # switch, to check the wiring

The bare form lists OBS's scenes, marks which are wired up, and warns about any scene = in your config that OBS does not have — a typo there is otherwise silent until you say the words.

OBS sources

A source command changes one source's visibility inside a scene rather than switching scenes:

[[commands]]
name = "ps5"
phrases = ["ps5", "ps 5", "playstation"]
source = "Cam Link Screen"
scene = "Main Screen"
visible = "toggle"          # show | hide | toggle (the default)
show_filter = "Move-In"     # Move filters on the scene
hide_filter = "Move-Out"
hide_delay_ms = 350

toggle is the useful default — "computa, ps5" puts it up and the same words take it down, so you never have to remember which way it was. Pair it with explicit visible = "show" / "hide" commands when you do want to say which way it goes.

Without a scene, the source is looked for in whichever scene is on program when you say it — right for a source that appears in several, and it means the same command works from wherever you are. Add scene = "Main Screen" to pin it to one scene instead; alongside source, scene says where the source lives and is not a scene switch. Sources inside a group are found without naming the group.

Animating it in and out

show_filter / hide_filter name Move filters on the scene, and turn the visibility flip into an animation. The order is the whole point, and it is not symmetric:

show:  make the source visible  ->  enable show_filter
hide:  enable hide_filter  ->  wait hide_delay_ms  ->  hide the source

A Move filter animates something that is on screen, so on the way in the source has to exist before it can move, and on the way out it has to survive until the move ends — hiding first would cut the animation off at frame one. That asymmetry, and the wait in the middle of the hide, is why this is a sequence rather than one request. It is the same sequence a Companion button runs for the same effect.

Enabling the filter is what triggers it; the plugin turns the filter back off when the move finishes, which is what lets the next command trigger it again. hide_delay_ms should be at least as long as the move or the source is yanked mid-animation — it defaults to 350, and shows up as about 350 ms of extra latency on a hide and none on a show.

A toggle needs both filters or neither: a move that is never undone parks the source wherever it left it — off screen, and invisible the next time it is shown. That is rejected at load. One-way show / hide commands only need their own filter.

voice-control obs sources                     # names, exactly as OBS has them
voice-control obs sources "Camera Screen"     # a scene other than the current one
voice-control obs filters "Main Screen"       # filter names, same deal
voice-control run "computa ps5"               # the whole sequence, to check the wiring
voice-control obs toggle "Cam Link Screen"    # bare flip, no filters

Copy names from sources and filters rather than from another tool's dropdown — OBS matches exactly, and it does not always spell a name the way something else displays it (Move-In, not Move-in).

sources marks which entries are currently hidden, which live in a group, and which are wired to a command — and warns about a source = the scene does not have, the same way the scene list does. A source that is in a different scene than the one you are on reads as "no source named …", with the names it does have listed; that message is the usual answer to a source command that does nothing.

It also warns when a scene holds two sources with the same name, which OBS allows. Which one a command gets is then OBS's answer to the name rather than ours — GetSceneItemId resolves it, so it is the same item any other obs-websocket client would move — but if the two are meant to behave differently, rename one.

Media keys

A media command presses one of the keyboard's transport keys:

[[commands]]
name = "next track"
phrases = ["next", "skip", "next song", "skip this song"]
media = "next"              # play_pause | next | previous
computa, skip        ->  media key: next track
computa, play        ->  media key: play/pause
computa, last song   ->  media key: previous track

macOS routes these to whichever app it considers to be playing, so this controls Spotify without talking to Spotify — and controls Music, or a browser tab, on the days one of those is making the noise instead. There is nothing about Spotify configured anywhere, and nothing to keep in sync when you switch players.

Spellings are interchangeable: play, pause, stop, resume and toggle all mean play_pause, skip means next, and prev / back mean previous.

There is one play/pause key, not a play and a pause. The hardware has only the one and every player treats it as a toggle, so "computa, pause" while already paused starts it playing again. Said once, either word does the obvious thing.

previous is the player's business, not ours: most read one press as "back to the start of this song" and only two as "the song before".

Accessibility

Synthesising a key press needs the Accessibility grant, which is separate from the microphone one and is never prompted for. Add ~/Library/Application Support/voice-control/voice-control under System Settings → Privacy & Security → Accessibility.

Without it macOS drops the event and says nothing, which looks exactly like a player that ignored it — so the grant is checked first and the command fails with a real message instead. The bare form is the quick way to tell that apart from a phrase that did not match:

voice-control media next          # no config read, no matching
voice-control run "computa skip"  # the command, matched and dispatched

A grant belongs to the responsible process, so run from a terminal the check reflects that terminal's Accessibility rather than the binary's. Only the launchd copy answers for itself.

Spotify controls

Use spotify when the command must target Spotify specifically or needs a control that has no media key. It implements every playback-changing Spotify Player API operation: play/resume, pause, next, previous, seek, repeat, shuffle, volume, queue, and Spotify Connect transfer. It also provides relative volume on top of the API.

Spotify requires a Premium account for Web API playback control. Create an app in the Spotify developer dashboard, select Web API, and register this exact redirect URI:

http://127.0.0.1:8888/callback

Put the app's public client id in the command file—there is no client secret because the CLI uses OAuth authorization code with PKCE:

[spotify]
client_id = "YOUR_CLIENT_ID"
redirect_uri = "http://127.0.0.1:8888/callback"
token_file = "~/.config/voice-control/spotify-refresh-token"

Then authorize it once. The browser asks only for playback read/write access, redirects to a short-lived loopback listener, and the CLI saves the refresh token with owner-only permissions:

voice-control spotify authorize
voice-control spotify devices
voice-control spotify volume 40

Access tokens are refreshed automatically. Spotify refresh tokens currently expire after six months, so rerun authorize when Spotify reports that the saved token has expired or been revoked. The client also persists a rotated refresh token if Spotify supplies one.

A Spotify action is an inline table on a command:

[[commands]]
name = "music louder"
phrases = ["music louder", "turn up the music", "spotify louder"]
spotify = { action = "volume_change", percent = 10 }

[[commands]]
name = "music quieter"
phrases = ["music quieter", "turn down the music", "spotify quieter"]
spotify = { action = "volume_change", percent = -10 }

[[commands]]
name = "music volume forty"
phrases = ["music volume forty", "spotify volume forty"]
spotify = { action = "volume", percent = 40 }

The complete action forms are:

spotify = { action = "play" }
spotify = { action = "pause" }
spotify = { action = "next" }
spotify = { action = "previous" }
spotify = { action = "seek", position_ms = 30000 }
spotify = { action = "repeat", state = "track" } # off | track | context
spotify = { action = "shuffle", state = true }
spotify = { action = "volume", percent = 40 }
spotify = { action = "volume_change", percent = -10 }
spotify = { action = "queue", uri = "spotify:track:..." }
spotify = { action = "transfer", device_id = "...", play = true }

All except transfer accept an optional device_id; without one they target the active device. Copy ids from voice-control spotify devices. Device ids are not guaranteed to remain stable, so omit one unless the command intentionally targets a particular Spotify Connect player.

play can resume with no other fields, start a context, or start an explicit list of tracks:

spotify = { action = "play", context_uri = "spotify:playlist:...", offset_position = 2, position_ms = 0 }
spotify = { action = "play", uris = ["spotify:track:...", "spotify:track:..."] }

offset_uri = "spotify:track:..." can replace offset_position for a context. A context and uris are mutually exclusive; invalid mixes and out-of-range volumes are rejected when the configuration loads.

Tones

sound plays a wav from SOUNDS_DIR, named without the extension:

[[commands]]
name = "ping"
phrases = ["ping", "hello", "hi", "are you listening", "you there"]
sound = "ping"

A command whose only step is one of these does nothing else, which is the whole point of it — something to say to check the daemon is awake and hearing you, that has no side effects if it turns out you were talking to yourself.

The generic success chirp is left off for a command that makes its own noise, so it answers once rather than twice. A failure still gets the failure tone: if the wav is missing the command has not done the one thing it was for, and it says so rather than sitting there silently.

ping.wav is two taps at one pitch. Every other cue is a pair of notes going somewhere — rising for wake and ok, falling for fail — so a flat double-tap is the one shape left that cannot be mistaken for any of them. scripts/make-sounds.py generates all four.

Audio devices

output and input move the system's default devices — the same thing as picking one in the Sound pane, done through the CoreAudio HAL directly. Devices are named once in a [devices] table and referred to by that name:

[devices]
speakers = "CalDigit USB-C Pro Audio"
headphones = "AirPods"
microphone = "Wireless microphone"

[[commands]]
name = "headphones"
phrases = ["headphones", "airpods", "air pods", "headset"]
output = "headphones"

[[commands]]
name = "speakers"
phrases = ["speakers", "speaker", "on speakers", "to speakers"]
output = "speakers"

Each value is a case-insensitive substring of the device name, because the HAL spells them out in full — "Dustin's AirPods Pro #3" — and the pairing renames itself often enough that matching the whole thing would be a config edit every few months. voice-control devices lists what those substrings are matched against, with the current defaults and which alias lands on which device:

$ voice-control devices
inputs
  Wireless microphone  (default, microphone)
  Cam Link 4K
  AirPods Pro 2  (headphones)
outputs
  CalDigit USB-C Pro Audio  (default, speakers)
  Mac mini Speakers
  AirPods Pro 2  (headphones)

The name must be in the table. An alias that is merely misspelt matches nothing, and so does a device that is only unplugged — there is no way to tell those apart at the point the command runs, so the spelling is checked at startup instead and a typo is a load error naming the aliases that do exist.

Alert sounds follow the output device, which is what the Sound pane does too. Anything that has pinned a device of its own does not: a call already running in Discord stays where it is, because that is macOS's rule and not this daemon's.

Switching the input does not move the daemon's own microphone. cpal opens a device, not "whatever is default", so the capture stream stays where it was until the daemon restarts. INPUT_DEVICE is what decides where it listens.

There is deliberately no headphones command for input, and adding one is a bad idea: selecting the AirPods microphone drops the whole bluetooth link into 16 kHz call mode, so the music in your ears gets worse in exchange for a microphone worse than the one already on the desk. stop-airpods-mic exists on this machine to undo exactly that when macOS does it unasked — and it will undo this too, within about 20 ms, which is the other reason not to.

The bare forms take an alias or, if it is not one, part of a device name — which is how you find out what to put in the table:

voice-control output headphones     # an alias from [devices]
voice-control output "Mac mini"     # or a substring, unconfigured

Asking for the device that is already default is a no-op and says so. That is not only cosmetic: writing the property fires the HAL notification it was set from, and there is another agent on this machine listening for that one.

Home Assistant

hass calls a service on one entity over Home Assistant's REST API. A plain url step cannot: the API wants a bearer token and a JSON body, and neither is something an HTTP step carries.

[hass]
url = "https://hass.lan"   # the base, no /api
token = ""                 # empty reads HASS_TOKEN instead
insecure = true            # accept the ingress certificate unchecked

[[commands]]
name = "speakers"
phrases = ["speakers", "speaker", "on speakers", "to speakers"]
hass = "switch.desk_speakers"
service = "turn_on"

service defaults to toggle, for the same reason an OBS source does — a plug command is usually said the same way twice. It is taken as belonging to the entity's own domain, so turn_on on a switch. entity is switch.turn_on. A service that names a domain of its own is used as written, which is how the domain-agnostic ones are reached:

service = "homeassistant.turn_off"   # works on anything

data is anything else the service takes, sent alongside the entity — data = { brightness = 128 }. A switch needs none of it.

Entities are enumerated one command at a time, the same as scenes: the daemon can only reach the ones you named. Copy the ids from:

voice-control hass                        # list them all
voice-control hass speaker                # filtered, on id or name
voice-control hass switch.desk_speakers   # toggle one, to check it
voice-control hass switch.desk_speakers turn_on

The listing marks which entities commands already use and warns about any hass = your Home Assistant does not have. That warning is the whole reason it exists: Home Assistant answers a call naming an entity it has never heard of with a perfectly happy 200 and no state change, so a wrong id is otherwise indistinguishable from a working command until the day you say the words and nothing happens. The shape of the id is checked at startup for the same reason, and the dispatch logs how many entities each call actually changed — a 0 there is either something already in that state or an id that does not exist.

insecure skips certificate verification entirely. A local ingress fronted by a CA only your network knows about is an ordinary way to run Home Assistant, and the alternative is teaching every machine about the CA. It does mean nothing is checking who answered, so leave it off for anything reached over the internet. It applies to the Home Assistant client alone — the HTTP steps get a client of their own, and a certificate exception for the plug in your office has no business weakening them.

Switching outputs by plug rather than by device belongs here: make "speakers" turn the plug on and "headphones" turn it off, in place of the output steps above. Both at once is a flow — plug on, then point the default output at it.

Programs

run executes a local program. Everything else a command can do reaches something over a socket — HTTP, obs-websocket, the HAL — which leaves anything that ships a CLI and no API out of reach:

[[commands]]
name = "desk lights"
phrases = ["lights", "desk lights", "toggle the lights"]
run = "~/bin/lights"
args = ["toggle"]
timeout_ms = 5000                    # default 10000
env = { LIGHTS_HOST = "10.0.0.4" }   # on top of the daemon's own

There is no shell in between. The program is executed directly, so one entry in args is one argument however many spaces are in it, and there is nothing to quote, glob or expand. Nothing a misheard phrase says can reach the arguments either: they come from the config file and never from the transcript.

Name the binary in full. launchd starts an agent with a PATH of /usr/bin:/bin:/usr/sbin:/sbin and nothing else, so a bare name that works in your shell — anything under /opt/homebrew/bin, say — is not on the daemon's PATH at all. A leading ~ is expanded here since no shell is running to do it. A program that is not where the config says is a warning at startup rather than a load error: it may only be one that is not installed yet, which is no reason for every other command in the file to stop working.

A non-zero exit fails the step — it is the only thing the program tells us, and chirping success over a program that just failed is worse than having no command at all. That stops the flow and plays the failure tone, with whatever it printed in the message. Output goes to the log on a successful run too, clipped to one line.

timeout_ms is when it gets killed. The dispatch holds the pipeline open while the program runs, so something that never comes back is the wake word not answering you until it does — and the things worth saying this out loud to (bluetooth, a sleeping device, a network hop) are exactly the ones that can hang.

voice-control run "computa lights" is the way to check one without saying anything, and prints the program and arguments it matched.

Flows

A command that does more than one thing is a list of steps, run in order, stopping at the first failure. A step takes the same fields a one-step command does, plus wait_ms for a step that only waits.

The reason the ps5 commands are flows: a Move filter animates the source into the scene it lives in, so saying "show ps5" from the camera scene has to get you there first.

[[commands]]
name = "show ps5"
phrases = ["show ps5", "show the ps5", "show playstation"]

  [[commands.steps]]
  scene = "Main Screen"

  [[commands.steps]]
  source = "Cam Link Screen"
  scene = "Main Screen"
  visible = "show"
  show_filter = "Move-In"
computa, show ps5   ->  obs scene "Main Screen"
                    ->  obs show source "Cam Link Screen" via Move-In

Everything above is the one-step shorthand for exactly this, so single-action commands need no [[commands.steps]] at all — and a command uses one form or the other, never both. Failures stop the flow rather than pressing on: if the scene switch did not happen there is no point enabling the move that was meant to play on it, and continuing would leave the source shown somewhere you cannot see it.

Steps are not limited to OBS — HTTP, media keys and programs are steps like any other, so one phrase can pause the music, mute Discord, switch scene and bring a source in:

[[commands]]
name = "brb"
phrases = ["brb", "be right back"]

  [[commands.steps]]
  media = "play_pause"

  [[commands.steps]]
  url = "{discord}/mute/on"

  [[commands.steps]]
  scene = "Main Screen"

  [[commands.steps]]
  wait_ms = 300          # let the scene transition land

  [[commands.steps]]
  source = "Showering text"
  scene = "Main Screen"
  visible = "show"

Errors name the step they came from (command "brb", step 5: …), which is the difference between a config you can fix and one you have to bisect. Check a whole flow without saying anything:

voice-control run "computa show ps5"

It prints the flow it matched before running it:

matched: show ps5 (score 1.000) -> obs scene "Main Screen" ->
  obs show source "Cam Link Screen" in "Main Screen" via Move-In

If you would rather not keep the password in the config file, leave it empty and set OBS_PASSWORD in the environment (the LaunchAgent reads it from .env via scripts/install-agent.sh). If you do put it in the file, chmod 600 ~/.config/voice-control/commands.toml.

The daemon opens a fresh connection per switch rather than holding one open. Commands are seconds apart at best, and a short-lived connection cannot rot when OBS restarts or the machine sleeps.

Menu bar

The daemon puts an icon in the menu bar. It is the answer to the two questions the logs used to be the only way to answer: is it hearing me at all, and what did it think I said?

mic          idle, waiting for the wake word
mic.fill     heard you, capturing the command
waveform     transcribing and dispatching
checkmark    the last command went through   (1.5s, then back to idle)
mic.slash    paused from the menu
!            not hearing anything - see below

They are SF Symbols rendered as template images, so they follow the menu bar in light, dark and tinted appearances. Hovering shows the status line without opening anything.

The menu holds the current state, the input device and a live level meter, how long ago it was last woken, and the last ten utterances with what became of each:

Listening for "computa"
Wireless microphone  ▃▅▂·····
Last woken 4m ago
────────────────────
Recent
   2m  "computa mute"    mute OK
   4m  "computa muted"   no match
────────────────────
Pause listening
Open logs
Restart agent
────────────────────
Quit until next login

The no-match lines are the point of the list. Growing the phrase lists in commands.toml used to mean grepping stdout.log; now the last ten are one click away, and the log is still there for the rest.

Pause listening stops the wake word from scoring without stopping the daemon - for a screen share, or a meeting where "computa" is going to come up. Audio keeps flowing, so the pre-roll stays warm and resuming is instant. A capture already in flight is allowed to finish.

Quit has to boot the job out of launchd, because KeepAlive would undo a plain exit within the second. It comes back at next login, or immediately with ./scripts/install-agent.sh.

Two distinct ways of hearing nothing get distinct warnings, because they have different causes:

Menu says Means
Silent for 30s - check microphone access audio is arriving, all of it empty: the TCC grant was revoked, or the mic is muted in hardware
No audio for 12s - the input device is gone no buffers at all: the device was unplugged, or CoreAudio dropped the stream

The second used to be invisible - the stream ending would exit the process, launchd would restart it, and nothing said so.

If no icon appears, check the log for menu bar item created. If it is there, the item exists and something is hiding it: Ice, Bartender and friends file new items into their hidden section by default, and it has to be dragged out (cmd-drag along the menu bar) once.

TRAY=false gives back the old headless daemon, which is what you want over ssh or under a debugger, where there is no window server to talk to. Every subcommand is headless regardless.

Status feed

Setting a URL under [status] makes the daemon post what it is doing to it. This exists for the OBS overlay, which lights an indicator when the wake word lands and shows what became of the command:

[status]
url = "http://127.0.0.1:8080/api/voice"
{
  "wake_word": "alexa",
  "state": "idle",
  "device": "Wireless microphone",
  "result": {
    "id": 42,
    "transcript": "next track",
    "outcome": "dispatched",
    "command": "next track"
  }
}

state is one of starting, idle, listening, thinking, paused, deaf, stalled, stopped — the same picture the menu bar draws, including the two fault states, which carry fault_ms. result is present only while the last utterance is recent, and its outcome is dispatched, failed, no_match or unheard. id counts utterances from startup, so that saying the same thing twice is distinguishable from one command being reported twice.

A post goes out when the picture changes and every five seconds regardless, which is what lets the far end tell a quiet daemon from a dead one. Nothing in the pipeline calls this — it watches the same Status the menu bar reads, which is also how the derived fault states get published without anything having to run a timer for them. An endpoint that is down is logged once and retried on the next change; it never blocks anything.

Environment

Variable Default Meaning
CONFIG_PATH ~/.config/voice-control/commands.toml command table
INPUT_DEVICE system default case-insensitive substring of the device name
SOUNDS_DIR unset (silent) directory holding wake.wav, ok.wav, fail.wav, and any wav a sound command names
OBS_PASSWORD unset obs-websocket password, if not in the config
HASS_TOKEN unset Home Assistant long-lived access token, if not in the config
SPOTIFY_CLIENT_ID unset Spotify application's public client id, if not in the config
SPOTIFY_REFRESH_TOKEN token file Spotify PKCE refresh token; primarily useful when secrets are injected
DSTN_LOG info tracing filter
TRAY true menu bar status item
LOG_DIR ~/Library/Application Support/voice-control/logs what the menu's "Open logs" opens
LAUNCHD_LABEL com.dstn.voice-control what the menu's restart and quit act on

Unmatched transcripts are logged at info on purpose — grepping them is how the phrase lists get grown:

grep "no matching command" ~/Library/Application\ Support/voice-control/logs/stdout.log

Notes

  • The main thread belongs to AppKit. NSStatusItem has to be created on it under a running NSApplication, so main is not #[tokio::main]: the runtime is built by hand and the pipeline is driven on a thread of its own. It gets a thread rather than a task because neither the wake word detector nor the VAD is Sync, so the pipeline future cannot go to a work-stealing scheduler. Everything it spawns - the resampler, whisper, HTTP - still lands on the runtime.
  • Speaker loopback. On speakers, someone in voice chat saying "computa mute" will trigger it. Headphones avoid this, and "computa, headphones" is now one of the things that can get you there — though not while they are the ones saying it.
  • Commands are a literal phrase → a fixed action. No parameters ("set volume to 30") without extending the matcher.
  • ort is pinned to =2.0.0-rc.10 in Cargo.toml, which is the version voice_activity_detector asks for. The wake word and the VAD both run on ONNX Runtime, and two ort versions in one binary means two copies of the runtime.

About

Computa, control this computer, thank you.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Contributors

Languages