A background macOS agent that listens for "computa, …" and turns the command that follows into an HTTP request, an OBS change, a media key, a Spotify Web API control, an audio device switch, a Home Assistant service call, or a local program.
computa, mute -> POST :8009/v1/voice/canary/mute/on
computa, unmute -> POST :8009/v1/voice/canary/mute/off
computa, deafen -> POST :8009/v1/voice/canary/deaf/on
computa, push to talk -> POST :8009/v1/voice/canary/mode/PUSH_TO_TALK
computa, desk mic -> POST :8080/v1/input/set/desk
computa, camera -> OBS scene "Camera Screen"
computa, main -> OBS scene "Main Screen"
computa, ps5 -> OBS scene "Main Screen", then animate the
"Cam Link Screen" source in (or out)
computa, hide ps5 -> animate "Cam Link Screen" out
computa, skip -> media key: next track
computa, play -> media key: play/pause
computa, music quieter -> Spotify: lower active device volume 10%
computa, headphones -> default output device: AirPods
computa, speakers -> default output device: the desk speakers
computa, coffee -> Home Assistant: switch.coffee_machine
toggle
computa, are you -> a tone back, and nothing else
listening?
Targets are discord-rpc-control on localhost and mic-api on the Linux box. Nothing is hardcoded — commands are a TOML table, so adding one is a config edit and a restart.
Discord RPC mutes inside Discord and leaves the OS input device open, which is what makes "computa, unmute" work while muted. Do not swap the dispatch for an OS-level mute; the daemon would go deaf to its own un-mute.
cpal capture ──► resample 48k→16k mono ──► 3s pre-roll ring
│
┌─────┴─────┐
│ IDLE │ openWakeWord scores every 80ms hop
└─────┬─────┘
"computa" ✔ → chirp
│
┌─────┴─────┐
│ LISTENING │ Silero VAD: up to 2.5s for the
└─────┬─────┘ command to start, then 700ms of
│ silence ends it
│ (seeded with 900ms of pre-roll)
│
whisper.cpp base.en → "computa mute"
│
normalise → suffix candidates → fuzzy match
│
commands.toml → POST … → ok/fail tone
Two stages rather than one because a wake-word model is cheap enough to run continuously and whisper is not.
The pre-roll is the part that is easy to get wrong. "computa, mute" is
one breath, so the command is already partly spoken when the wake word
starts scoring — and the detector then needs patience more hops
before it commits. Between them the whole utterance can be in the past
by the time the detector says anything, so the window reaches back far
enough to recover the wake word too. That the wake word ends up in the
clip is fine: the matcher works over suffixes and drops it.
It also means the capture starts out having already heard speech —
the wake word — whatever the user does next. Ending it on trailing
silence alone therefore closes the window about 700 ms after "computa",
which is not long enough to say "computa", think for a beat, and then
give the command: the first syllable arrives after the capture has
already been sent to whisper. So silence does not end anything until
something has been said live, past the end of the pre-roll. Until
then the only clock running is grace_ms, and when that runs out the
wake word is written off as a false trigger and dropped quietly.
Nothing is lost in the brisk case. Said in one breath, "computa mute" is still being spoken when the detector commits, so its tail arrives live, and the hangover ends the capture as promptly as it ever did — the grace period only applies to a capture where nothing has been said yet, and there is nothing to cut short.
[listen]
grace_ms = 2500 # how long the command may still be starting
silence_ms = 700 # trailing silence that ends one already started
max_ms = 8000 # ceiling on a capture, pre-roll includedBoth stages sit behind traits (wake::Detector, stt::Transcriber),
and the wake word is openWakeWord through that trait, so the
engine is one file: src/wake/oww.rs.
It is three ONNX graphs in series, scored once per 80 ms hop:
1280 samples ─► melspectrogram ─► 8 mel frames of 32 bins
│ (ring of 80)
last 76 frames ▼
embedding ─► 96 floats
│ (ring of 16)
▼
wake ─► one probability
The first two are openWakeWord's own pretrained feature extractors and are the same whatever the wake word is; only the third is trained per word. That split is the whole reason this replaced rustpotter, which matched a waveform against a dozen recordings of it and could not tell the wake word from a television. The embedding model is a speech representation trained on a very large corpus, so the classifier is asked "what was said", not "how close is this to my takes" — a working model reads ~0.00 on speech that is not its word rather than 0.6.
Inference runs on ONNX Runtime via ort, which was already in the tree:
the Silero VAD uses it too.
Needs cmake (brew install cmake) to build whisper.cpp. The first
build takes a few minutes; that is whisper.cpp compiling, not a hang.
mkdir -p ~/.config/voice-control
curl -L -o ~/.config/voice-control/ggml-base.en-q5_1.bin \
https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en-q5_1.binbase.en is noticeably better than tiny.en on one- and two-word
commands and still runs in well under 200 ms on Apple Silicon. Drop to
ggml-tiny.en-q5_1.bin if you want the latency back.
Three files: the two shared feature extractors, which never change, and a classifier for the word itself.
cd ~/.config/voice-control
base=https://github.com/dscripka/openWakeWord/releases/download/v0.5.1
curl -L -O $base/melspectrogram.onnx
curl -L -O $base/embedding_model.onnx
curl -L -O $base/alexa_v0.1.onnx && mv alexa_v0.1.onnx alexa.onnxalexa, hey_jarvis, hey_mycroft and hey_rhasspy are the
pretrained words, and they are very good: alexa_v0.1 reads 1.000 on
"alexa" and below 0.07 on everything else tried here, including real
speech, "computer", and a directory of takes of a different wake word.
Point [wake] model at whichever one you want.
Any other word has to be trained, which openWakeWord does with a corpus of synthetic speech, not a handful of takes — see its training notebooks. A model trained on too few negatives is worse than useless here: one tried during this port fired at 0.98+ on ordinary dialogue that sounded nothing like its word, which is exactly the failure that made rustpotter unusable. Check any model you are given before trusting it:
# Should all fire.
cargo run --release -- score positives/*.wav
# Should all read ~0.00. If they do not, the model is not usable,
# whatever it scores on the word itself.
cargo run --release -- score negatives/*.wavcp commands.example.toml ~/.config/voice-control/commands.tomlSet mic under [targets] to the Linux host, and check the Discord
branch in the discord URL matches DISCORD_BRANCH in
discord-rpc-control (its installer defaults to canary).
cargo run --release -- listenPrints an input level once a second and a line per detection:
level 0.067 ###
1 alexa score 0.998
If the level stays at zero, the microphone permission is the problem, not the model — it says so after five seconds of silence.
With a well-trained model there is not much to tune: scores sit at either end of the range, so anything from 0.3 to 0.8 behaves the same and 0.5 is fine. Tuning only starts to matter when the model is marginal, and then the honest answer is usually a better model.
Guessing at it from a live microphone is slow, so there are two
subcommands for doing it from recordings instead. record writes the
microphone to a directory of one-minute wavs and prints every hit as it
happens, with the file and offset:
cargo run --release -- record ~/negatives # leave it running
1 audio-0003.wav 12.4s score 0.981 # ← go and listen to thatLeave it going through whatever sets the daemon off. Then score
replays wavs through the detector — no microphone, no whisper, nothing
dispatched — and prints the peak score, the longest run of hops over
the threshold, and whether it would have fired:
cargo run --release -- score ~/negatives/*.wav
audio-0003.wav peak 0.981 at 12.4s run 6 FIRED at 12.2sthreshold wants to sit in the gap between what your voice scores and
what the room scores. If there is no gap — if run is as long on
television as it is on you — no threshold will fix it.
patience is the other lever: how many consecutive 80 ms hops have to
clear the threshold before it counts. It costs patience × 80 ms of
latency, which the pre-roll absorbs, and it drops false positives that
only glance off the model for a frame or two. Two is a good default;
raising it trades a little responsiveness for fewer of them.
./scripts/install-agent.shBuilds release, installs to ~/Library/Application Support/voice-control,
ad-hoc signs the binary (so the microphone grant survives rebuilds), and
bootstraps the LaunchAgent.
macOS will ask for microphone access the first time it runs. If no prompt appears and it never hears anything, check System Settings → Privacy & Security → Microphone.
Media key commands need a second grant, Accessibility, which is
never prompted for — add
~/Library/Application Support/voice-control/voice-control under
System Settings → Privacy & Security → Accessibility by hand. Skip it
if you have no media commands. Both grants are keyed to the code
signature, which is why the script ad-hoc signs the binary before
launchd sees it: without that, every rebuild would look like a brand
new app and drop them.
voice-control run the daemon
voice-control devices list audio devices and their aliases
voice-control output headphones switch the default output, bare
voice-control input "Wireless" switch the default input, bare
voice-control listen wake word detections only
voice-control transcribe f.wav STT + matcher against a file
voice-control run "computa ps5" match a phrase and dispatch it
voice-control replay f.wav whole pipeline against a file
voice-control hass list home assistant entities
voice-control hass speaker the ones matching a filter
voice-control hass switch.desk_speakers call a service, bare
voice-control obs list obs scenes
voice-control obs "Camera Screen" switch to a scene
voice-control obs sources list the sources in the current scene
voice-control obs filters "Main Screen" list a scene's filters
voice-control obs toggle "Cam Link Screen" flip a source, bare
voice-control media next press a media key, bare
voice-control spotify authorize authorize Spotify once with PKCE
voice-control spotify devices list Spotify Connect device ids
voice-control spotify volume 40 set Spotify's volume, bare
transcribe is the fast way to grow phrase lists without talking to a
microphone. run takes the phrase directly and dispatches it exactly
as the daemon would, sequences and waits included — the way to check
what a command does once you know it matches. replay runs a file
through wake word, VAD, STT, matching and dispatch — everything except
opening the microphone — so you can check a change end to end. Generate
test clips with say -v Samantha -o x.aiff "computa mute" piped
through afconvert -f WAVE -d LEI16@16000 -c 1 x.aiff x.wav.
Scene switching goes over obs-websocket 5.x. Enable the server in OBS under Tools -> WebSocket Server Settings, then:
[obs]
url = "ws://127.0.0.1:4455" # wss:// for an OBS on another machine
password = "" # empty reads OBS_PASSWORD instead
[[commands]]
name = "camera scene"
phrases = ["camera", "cam", "show me"]
scene = "Camera Screen"Scenes are enumerated one command at a time on purpose: the daemon can
only switch to scenes you named, never to whatever it thought it heard.
scene is matched by OBS exactly, so copy the names from:
voice-control obs # list them
voice-control obs "Camera Screen" # switch, to check the wiringThe bare form lists OBS's scenes, marks which are wired up, and warns about any
scene = in your config that OBS does not have — a typo there is
otherwise silent until you say the words.
A source command changes one source's visibility inside a scene
rather than switching scenes:
[[commands]]
name = "ps5"
phrases = ["ps5", "ps 5", "playstation"]
source = "Cam Link Screen"
scene = "Main Screen"
visible = "toggle" # show | hide | toggle (the default)
show_filter = "Move-In" # Move filters on the scene
hide_filter = "Move-Out"
hide_delay_ms = 350toggle is the useful default — "computa, ps5" puts it up and the same
words take it down, so you never have to remember which way it was.
Pair it with explicit visible = "show" / "hide" commands when you do
want to say which way it goes.
Without a scene, the source is looked for in whichever scene is on
program when you say it — right for a source that appears in several,
and it means the same command works from wherever you are. Add
scene = "Main Screen" to pin it to one scene instead; alongside
source, scene says where the source lives and is not a scene
switch. Sources inside a group are found without naming the group.
show_filter / hide_filter name Move filters on the scene,
and turn the visibility flip into an animation. The order is the whole
point, and it is not symmetric:
show: make the source visible -> enable show_filter
hide: enable hide_filter -> wait hide_delay_ms -> hide the source
A Move filter animates something that is on screen, so on the way in the source has to exist before it can move, and on the way out it has to survive until the move ends — hiding first would cut the animation off at frame one. That asymmetry, and the wait in the middle of the hide, is why this is a sequence rather than one request. It is the same sequence a Companion button runs for the same effect.
Enabling the filter is what triggers it; the plugin turns the filter
back off when the move finishes, which is what lets the next command
trigger it again. hide_delay_ms should be at least as long as the
move or the source is yanked mid-animation — it defaults to 350, and
shows up as about 350 ms of extra latency on a hide and none on a show.
A toggle needs both filters or neither: a move that is never undone
parks the source wherever it left it — off screen, and invisible the
next time it is shown. That is rejected at load. One-way show /
hide commands only need their own filter.
voice-control obs sources # names, exactly as OBS has them
voice-control obs sources "Camera Screen" # a scene other than the current one
voice-control obs filters "Main Screen" # filter names, same deal
voice-control run "computa ps5" # the whole sequence, to check the wiring
voice-control obs toggle "Cam Link Screen" # bare flip, no filtersCopy names from sources and filters rather than from another tool's
dropdown — OBS matches exactly, and it does not always spell a name the
way something else displays it (Move-In, not Move-in).
sources marks which entries are currently hidden, which live in a
group, and which are wired to a command — and warns about a source =
the scene does not have, the same way the scene list does. A source
that is in a different scene than the one you are on reads as "no
source named …", with the names it does have listed; that message is
the usual answer to a source command that does nothing.
It also warns when a scene holds two sources with the same name,
which OBS allows. Which one a command gets is then OBS's answer to the
name rather than ours — GetSceneItemId resolves it, so it is the same
item any other obs-websocket client would move — but if the two are
meant to behave differently, rename one.
A media command presses one of the keyboard's transport keys:
[[commands]]
name = "next track"
phrases = ["next", "skip", "next song", "skip this song"]
media = "next" # play_pause | next | previouscomputa, skip -> media key: next track
computa, play -> media key: play/pause
computa, last song -> media key: previous track
macOS routes these to whichever app it considers to be playing, so this controls Spotify without talking to Spotify — and controls Music, or a browser tab, on the days one of those is making the noise instead. There is nothing about Spotify configured anywhere, and nothing to keep in sync when you switch players.
Spellings are interchangeable: play, pause, stop, resume and
toggle all mean play_pause, skip means next, and prev / back
mean previous.
There is one play/pause key, not a play and a pause. The hardware has only the one and every player treats it as a toggle, so "computa, pause" while already paused starts it playing again. Said once, either word does the obvious thing.
previous is the player's business, not ours: most read one press as
"back to the start of this song" and only two as "the song before".
Synthesising a key press needs the Accessibility grant, which is
separate from the microphone one and is never prompted for. Add
~/Library/Application Support/voice-control/voice-control under
System Settings → Privacy & Security → Accessibility.
Without it macOS drops the event and says nothing, which looks exactly like a player that ignored it — so the grant is checked first and the command fails with a real message instead. The bare form is the quick way to tell that apart from a phrase that did not match:
voice-control media next # no config read, no matching
voice-control run "computa skip" # the command, matched and dispatchedA grant belongs to the responsible process, so run from a terminal the check reflects that terminal's Accessibility rather than the binary's. Only the launchd copy answers for itself.
Use spotify when the command must target Spotify specifically or
needs a control that has no media key. It implements every
playback-changing Spotify Player API operation: play/resume, pause,
next, previous, seek, repeat, shuffle, volume, queue, and Spotify
Connect transfer. It also provides relative volume on top of the API.
Spotify requires a Premium account for Web API playback control. Create an app in the Spotify developer dashboard, select Web API, and register this exact redirect URI:
http://127.0.0.1:8888/callback
Put the app's public client id in the command file—there is no client secret because the CLI uses OAuth authorization code with PKCE:
[spotify]
client_id = "YOUR_CLIENT_ID"
redirect_uri = "http://127.0.0.1:8888/callback"
token_file = "~/.config/voice-control/spotify-refresh-token"Then authorize it once. The browser asks only for playback read/write access, redirects to a short-lived loopback listener, and the CLI saves the refresh token with owner-only permissions:
voice-control spotify authorize
voice-control spotify devices
voice-control spotify volume 40Access tokens are refreshed automatically. Spotify refresh tokens
currently expire after six months, so rerun authorize when Spotify
reports that the saved token has expired or been revoked. The client
also persists a rotated refresh token if Spotify supplies one.
A Spotify action is an inline table on a command:
[[commands]]
name = "music louder"
phrases = ["music louder", "turn up the music", "spotify louder"]
spotify = { action = "volume_change", percent = 10 }
[[commands]]
name = "music quieter"
phrases = ["music quieter", "turn down the music", "spotify quieter"]
spotify = { action = "volume_change", percent = -10 }
[[commands]]
name = "music volume forty"
phrases = ["music volume forty", "spotify volume forty"]
spotify = { action = "volume", percent = 40 }The complete action forms are:
spotify = { action = "play" }
spotify = { action = "pause" }
spotify = { action = "next" }
spotify = { action = "previous" }
spotify = { action = "seek", position_ms = 30000 }
spotify = { action = "repeat", state = "track" } # off | track | context
spotify = { action = "shuffle", state = true }
spotify = { action = "volume", percent = 40 }
spotify = { action = "volume_change", percent = -10 }
spotify = { action = "queue", uri = "spotify:track:..." }
spotify = { action = "transfer", device_id = "...", play = true }All except transfer accept an optional device_id; without one they
target the active device. Copy ids from voice-control spotify devices.
Device ids are not guaranteed to remain stable, so omit one unless the
command intentionally targets a particular Spotify Connect player.
play can resume with no other fields, start a context, or start an
explicit list of tracks:
spotify = { action = "play", context_uri = "spotify:playlist:...", offset_position = 2, position_ms = 0 }
spotify = { action = "play", uris = ["spotify:track:...", "spotify:track:..."] }offset_uri = "spotify:track:..." can replace offset_position for a
context. A context and uris are mutually exclusive; invalid mixes and
out-of-range volumes are rejected when the configuration loads.
sound plays a wav from SOUNDS_DIR, named without the extension:
[[commands]]
name = "ping"
phrases = ["ping", "hello", "hi", "are you listening", "you there"]
sound = "ping"A command whose only step is one of these does nothing else, which is the whole point of it — something to say to check the daemon is awake and hearing you, that has no side effects if it turns out you were talking to yourself.
The generic success chirp is left off for a command that makes its own noise, so it answers once rather than twice. A failure still gets the failure tone: if the wav is missing the command has not done the one thing it was for, and it says so rather than sitting there silently.
ping.wav is two taps at one pitch. Every other cue is a pair of notes
going somewhere — rising for wake and ok, falling for fail — so a flat
double-tap is the one shape left that cannot be mistaken for any of
them. scripts/make-sounds.py generates all four.
output and input move the system's default devices — the same thing
as picking one in the Sound pane, done through the CoreAudio HAL
directly. Devices are named once in a [devices] table and referred to
by that name:
[devices]
speakers = "CalDigit USB-C Pro Audio"
headphones = "AirPods"
microphone = "Wireless microphone"
[[commands]]
name = "headphones"
phrases = ["headphones", "airpods", "air pods", "headset"]
output = "headphones"
[[commands]]
name = "speakers"
phrases = ["speakers", "speaker", "on speakers", "to speakers"]
output = "speakers"Each value is a case-insensitive substring of the device name,
because the HAL spells them out in full — "Dustin's AirPods Pro #3" —
and the pairing renames itself often enough that matching the whole
thing would be a config edit every few months. voice-control devices
lists what those substrings are matched against, with the current
defaults and which alias lands on which device:
$ voice-control devices
inputs
Wireless microphone (default, microphone)
Cam Link 4K
AirPods Pro 2 (headphones)
outputs
CalDigit USB-C Pro Audio (default, speakers)
Mac mini Speakers
AirPods Pro 2 (headphones)
The name must be in the table. An alias that is merely misspelt matches nothing, and so does a device that is only unplugged — there is no way to tell those apart at the point the command runs, so the spelling is checked at startup instead and a typo is a load error naming the aliases that do exist.
Alert sounds follow the output device, which is what the Sound pane does too. Anything that has pinned a device of its own does not: a call already running in Discord stays where it is, because that is macOS's rule and not this daemon's.
Switching the input does not move the daemon's own microphone. cpal
opens a device, not "whatever is default", so the capture stream stays
where it was until the daemon restarts. INPUT_DEVICE is what decides
where it listens.
There is deliberately no headphones command for input, and adding one
is a bad idea: selecting the AirPods microphone drops the whole
bluetooth link into 16 kHz call mode, so the music in your ears gets
worse in exchange for a microphone worse than the one already on the
desk. stop-airpods-mic exists on this machine to undo exactly
that when macOS does it unasked — and it will undo this too, within
about 20 ms, which is the other reason not to.
The bare forms take an alias or, if it is not one, part of a device name — which is how you find out what to put in the table:
voice-control output headphones # an alias from [devices]
voice-control output "Mac mini" # or a substring, unconfiguredAsking for the device that is already default is a no-op and says so. That is not only cosmetic: writing the property fires the HAL notification it was set from, and there is another agent on this machine listening for that one.
hass calls a service on one entity over Home Assistant's REST API. A
plain url step cannot: the API wants a bearer token and a JSON body,
and neither is something an HTTP step carries.
[hass]
url = "https://hass.lan" # the base, no /api
token = "" # empty reads HASS_TOKEN instead
insecure = true # accept the ingress certificate unchecked
[[commands]]
name = "speakers"
phrases = ["speakers", "speaker", "on speakers", "to speakers"]
hass = "switch.desk_speakers"
service = "turn_on"service defaults to toggle, for the same reason an OBS source
does — a plug command is usually said the same way twice. It is taken
as belonging to the entity's own domain, so turn_on on a switch.
entity is switch.turn_on. A service that names a domain of its own is
used as written, which is how the domain-agnostic ones are reached:
service = "homeassistant.turn_off" # works on anythingdata is anything else the service takes, sent alongside the entity —
data = { brightness = 128 }. A switch needs none of it.
Entities are enumerated one command at a time, the same as scenes: the daemon can only reach the ones you named. Copy the ids from:
voice-control hass # list them all
voice-control hass speaker # filtered, on id or name
voice-control hass switch.desk_speakers # toggle one, to check it
voice-control hass switch.desk_speakers turn_onThe listing marks which entities commands already use and warns about
any hass = your Home Assistant does not have. That warning is the
whole reason it exists: Home Assistant answers a call naming an
entity it has never heard of with a perfectly happy 200 and no state
change, so a wrong id is otherwise indistinguishable from a working
command until the day you say the words and nothing happens. The
shape of the id is checked at startup for the same reason, and the
dispatch logs how many entities each call actually changed — a 0
there is either something already in that state or an id that does not
exist.
insecure skips certificate verification entirely. A local ingress
fronted by a CA only your network knows about is an ordinary way to run
Home Assistant, and the alternative is teaching every machine about the
CA. It does mean nothing is checking who answered, so leave it off for
anything reached over the internet. It applies to the Home Assistant
client alone — the HTTP steps get a client of their own, and a
certificate exception for the plug in your office has no business
weakening them.
Switching outputs by plug rather than by device belongs here: make
"speakers" turn the plug on and "headphones" turn it off, in place of
the output steps above. Both at once is a flow — plug on,
then point the default output at it.
run executes a local program. Everything else a command can do
reaches something over a socket — HTTP, obs-websocket, the HAL — which
leaves anything that ships a CLI and no API out of reach:
[[commands]]
name = "desk lights"
phrases = ["lights", "desk lights", "toggle the lights"]
run = "~/bin/lights"
args = ["toggle"]
timeout_ms = 5000 # default 10000
env = { LIGHTS_HOST = "10.0.0.4" } # on top of the daemon's ownThere is no shell in between. The program is executed directly, so
one entry in args is one argument however many spaces are in it, and
there is nothing to quote, glob or expand. Nothing a misheard phrase
says can reach the arguments either: they come from the config file and
never from the transcript.
Name the binary in full. launchd starts an agent with a PATH of
/usr/bin:/bin:/usr/sbin:/sbin and nothing else, so a bare name that
works in your shell — anything under /opt/homebrew/bin, say — is not
on the daemon's PATH at all. A leading ~ is expanded here since no
shell is running to do it. A program that is not where the config says
is a warning at startup rather than a load error: it may only be one
that is not installed yet, which is no reason for every other command
in the file to stop working.
A non-zero exit fails the step — it is the only thing the program tells us, and chirping success over a program that just failed is worse than having no command at all. That stops the flow and plays the failure tone, with whatever it printed in the message. Output goes to the log on a successful run too, clipped to one line.
timeout_ms is when it gets killed. The dispatch holds the pipeline
open while the program runs, so something that never comes back is the
wake word not answering you until it does — and the things worth saying
this out loud to (bluetooth, a sleeping device, a network hop) are
exactly the ones that can hang.
voice-control run "computa lights" is the way to check one without
saying anything, and prints the program and arguments it matched.
A command that does more than one thing is a list of steps, run in
order, stopping at the first failure. A step takes the same fields a
one-step command does, plus wait_ms for a step that only waits.
The reason the ps5 commands are flows: a Move filter animates the source into the scene it lives in, so saying "show ps5" from the camera scene has to get you there first.
[[commands]]
name = "show ps5"
phrases = ["show ps5", "show the ps5", "show playstation"]
[[commands.steps]]
scene = "Main Screen"
[[commands.steps]]
source = "Cam Link Screen"
scene = "Main Screen"
visible = "show"
show_filter = "Move-In"computa, show ps5 -> obs scene "Main Screen"
-> obs show source "Cam Link Screen" via Move-In
Everything above is the one-step shorthand for exactly this, so
single-action commands need no [[commands.steps]] at all — and a
command uses one form or the other, never both. Failures stop the flow
rather than pressing on: if the scene switch did not happen there is no
point enabling the move that was meant to play on it, and continuing
would leave the source shown somewhere you cannot see it.
Steps are not limited to OBS — HTTP, media keys and programs are steps like any other, so one phrase can pause the music, mute Discord, switch scene and bring a source in:
[[commands]]
name = "brb"
phrases = ["brb", "be right back"]
[[commands.steps]]
media = "play_pause"
[[commands.steps]]
url = "{discord}/mute/on"
[[commands.steps]]
scene = "Main Screen"
[[commands.steps]]
wait_ms = 300 # let the scene transition land
[[commands.steps]]
source = "Showering text"
scene = "Main Screen"
visible = "show"Errors name the step they came from (command "brb", step 5: …), which
is the difference between a config you can fix and one you have to
bisect. Check a whole flow without saying anything:
voice-control run "computa show ps5"It prints the flow it matched before running it:
matched: show ps5 (score 1.000) -> obs scene "Main Screen" ->
obs show source "Cam Link Screen" in "Main Screen" via Move-In
If you would rather not keep the password in the config file, leave it
empty and set OBS_PASSWORD in the environment (the LaunchAgent reads
it from .env via scripts/install-agent.sh). If you do put it in the
file, chmod 600 ~/.config/voice-control/commands.toml.
The daemon opens a fresh connection per switch rather than holding one open. Commands are seconds apart at best, and a short-lived connection cannot rot when OBS restarts or the machine sleeps.
The daemon puts an icon in the menu bar. It is the answer to the two questions the logs used to be the only way to answer: is it hearing me at all, and what did it think I said?
mic idle, waiting for the wake word
mic.fill heard you, capturing the command
waveform transcribing and dispatching
checkmark the last command went through (1.5s, then back to idle)
mic.slash paused from the menu
! not hearing anything - see below
They are SF Symbols rendered as template images, so they follow the menu bar in light, dark and tinted appearances. Hovering shows the status line without opening anything.
The menu holds the current state, the input device and a live level meter, how long ago it was last woken, and the last ten utterances with what became of each:
Listening for "computa"
Wireless microphone ▃▅▂·····
Last woken 4m ago
────────────────────
Recent
2m "computa mute" mute OK
4m "computa muted" no match
────────────────────
Pause listening
Open logs
Restart agent
────────────────────
Quit until next login
The no-match lines are the point of the list. Growing the phrase lists
in commands.toml used to mean grepping stdout.log; now the last ten
are one click away, and the log is still there for the rest.
Pause listening stops the wake word from scoring without stopping the daemon - for a screen share, or a meeting where "computa" is going to come up. Audio keeps flowing, so the pre-roll stays warm and resuming is instant. A capture already in flight is allowed to finish.
Quit has to boot the job out of launchd, because KeepAlive would
undo a plain exit within the second. It comes back at next login, or
immediately with ./scripts/install-agent.sh.
Two distinct ways of hearing nothing get distinct warnings, because they have different causes:
| Menu says | Means |
|---|---|
Silent for 30s - check microphone access |
audio is arriving, all of it empty: the TCC grant was revoked, or the mic is muted in hardware |
No audio for 12s - the input device is gone |
no buffers at all: the device was unplugged, or CoreAudio dropped the stream |
The second used to be invisible - the stream ending would exit the process, launchd would restart it, and nothing said so.
If no icon appears, check the log for menu bar item created. If it is
there, the item exists and something is hiding it: Ice, Bartender and
friends file new items into their hidden section by default, and it has
to be dragged out (cmd-drag along the menu bar) once.
TRAY=false gives back the old headless daemon, which is what you want
over ssh or under a debugger, where there is no window server to talk
to. Every subcommand is headless regardless.
Setting a URL under [status] makes the daemon post what it is doing
to it. This exists for the OBS overlay, which lights an indicator when
the wake word lands and shows what became of the command:
[status]
url = "http://127.0.0.1:8080/api/voice"{
"wake_word": "alexa",
"state": "idle",
"device": "Wireless microphone",
"result": {
"id": 42,
"transcript": "next track",
"outcome": "dispatched",
"command": "next track"
}
}state is one of starting, idle, listening, thinking,
paused, deaf, stalled, stopped — the same picture the menu bar
draws, including the two fault states, which carry fault_ms. result
is present only while the last utterance is recent, and its outcome
is dispatched, failed, no_match or unheard. id counts
utterances from startup, so that saying the same thing twice is
distinguishable from one command being reported twice.
A post goes out when the picture changes and every five seconds
regardless, which is what lets the far end tell a quiet daemon from a
dead one. Nothing in the pipeline calls this — it watches the same
Status the menu bar reads, which is also how the derived fault states
get published without anything having to run a timer for them. An
endpoint that is down is logged once and retried on the next change; it
never blocks anything.
| Variable | Default | Meaning |
|---|---|---|
CONFIG_PATH |
~/.config/voice-control/commands.toml |
command table |
INPUT_DEVICE |
system default | case-insensitive substring of the device name |
SOUNDS_DIR |
unset (silent) | directory holding wake.wav, ok.wav, fail.wav, and any wav a sound command names |
OBS_PASSWORD |
unset | obs-websocket password, if not in the config |
HASS_TOKEN |
unset | Home Assistant long-lived access token, if not in the config |
SPOTIFY_CLIENT_ID |
unset | Spotify application's public client id, if not in the config |
SPOTIFY_REFRESH_TOKEN |
token file | Spotify PKCE refresh token; primarily useful when secrets are injected |
DSTN_LOG |
info |
tracing filter |
TRAY |
true |
menu bar status item |
LOG_DIR |
~/Library/Application Support/voice-control/logs |
what the menu's "Open logs" opens |
LAUNCHD_LABEL |
com.dstn.voice-control |
what the menu's restart and quit act on |
Unmatched transcripts are logged at info on purpose — grepping them is
how the phrase lists get grown:
grep "no matching command" ~/Library/Application\ Support/voice-control/logs/stdout.log- The main thread belongs to AppKit.
NSStatusItemhas to be created on it under a runningNSApplication, somainis not#[tokio::main]: the runtime is built by hand and the pipeline is driven on a thread of its own. It gets a thread rather than a task because neither the wake word detector nor the VAD isSync, so the pipeline future cannot go to a work-stealing scheduler. Everything it spawns - the resampler, whisper, HTTP - still lands on the runtime. - Speaker loopback. On speakers, someone in voice chat saying "computa mute" will trigger it. Headphones avoid this, and "computa, headphones" is now one of the things that can get you there — though not while they are the ones saying it.
- Commands are a literal phrase → a fixed action. No parameters ("set volume to 30") without extending the matcher.
ortis pinned to=2.0.0-rc.10inCargo.toml, which is the versionvoice_activity_detectorasks for. The wake word and the VAD both run on ONNX Runtime, and twoortversions in one binary means two copies of the runtime.