Notes: design, measurements, gotchas
Notes: design, measurements, gotchas
Written for someone adapting this to their own machine. README.md is the how-to; this is the why, plus every trap we fell into so you do not have to. Numbers are from the author's laptop (8 cores, CPU only) unless noted.
1. Architecture
mic ── pw-record (16 kHz mono s16) ──► Capture ── 30 ms frames, RMS energy VAD, adaptive noise floor
│ cuts a phrase at ~0.7 s of silence (or at 14 s, at the quietest point)
▼
worker thread ── Phonon-2 decode (CPU, one phrase at a time)
▼
mode at the moment of the cut decides what a phrase means:
dictating ► typing helper ("new line"…) or wtype the text listening ► ignore unless it starts with
"computer": phrase table → user-commands.json
Hyprland keybinds ─► phonon-ctl.py ─► unix socket ─► daemon → Claude session → (spoken confirm) → run
run() ─► result: file → viewer on special workspace │ short text → notification │ none → "Done"
daemon writes 0/1/2 to $XDG_RUNTIME_DIR/phonon-recording ─► status-bar indicator
One continuous capture serves both modes; listening and dictation can be on together (dictation wins while on).
Everything is one Python process plus the pw-record child; there is no systemd unit (the author autostarts it from
Hyprland's hyprland.start).
2. The speech model
- Phonon-2 ships as
phonon-2.bps.tar.zst(164 MB:model.fermion,config.jsonwith the vocabulary, manifests).model.fermionis a packed container (five-value + int6 quantisation, ~2 bits per encoder weight). - Engine (since 2026-10-01): Fermion's own CPU engine from the
fermion-researchPyPI package, loaded straight frommodel/byfermion._speech.engine_phonon2_cpu.load(see_load_fermionin the daemon). On Linux x86-64 it uses a packed int8 encoder + a C TDT decode loop from the package's bundled.so(tier chosen from the CPU: AVX2 here; AVX-512 VNNI / AMX parts are faster). First load unpacks the encoder planes intomodel/cpu_planes_v1.bin(304 MB, ~30 s once), later loads take ~6-10 s. Measured on this laptop (i5-10210U, 4 threads): 1.5 GB resident (the fp32 path: 2.9 GB), 11-12x realtime (fp32: 8-10x), transcripts identical on two clips. Their 46x-157x figures are for AVX-512 Zen 5 machines; here it is a memory/disk win, not a latency win. - Our earlier claim was wrong: this file used to say the packed runtime was Apple-silicon only and that we had to
expand to fp32. The package docstring (
engine_phonon2_cpu.pyheader) still describes the older fp32-only tier, which is what misled us; the code below it (mode = FERMION_P2_CPU or "onedot") picks the packed engine.FERMION_P2_CPU=fp32forces the dense path. - Gotcha: do not put
vendor/onsys.pathbefore importingfermion. The package refuses to load if a flat module calledfermion_containeris already bound to a different file (ImportError, which the daemon then reports and survives by falling back). The vendored reader is therefore imported lazily inside the fallback only. - Fallback:
PHONON_ENGINE=hf(also used automatically if the Fermion engine fails to load) is the earlier implementation:vendor/fermion_container.pyexpands the container to fp32 and loads it into stock Hugging FaceParakeetForTDT, cached ashf/(2.4 GB safetensors, built on first use, ~30 s; ~5 s afterwards), ~3 GB resident, ~7x realtime. Verified word-for-word identical before/after that cache. Do not build that model on themetadevice:load_state_dict(..., assign=True)leaves the non-persistent positional buffers on meta ("cannot copy out of meta tensor"); build on CPU, then assign. Naive building held ~6.8 GB;assign=True, deleting raw tensors andmalloc_trimfixed it. - Threads: the daemon passes
PHONON_THREADS(4) asFERMION_CPU_THREADS(the engine's own default here would be 8, one per logical CPU) so the desktop keeps some CPU. - Their quoted 5.21 % average WER on English leaderboards is their number; we only spot-checked.
- English only. One issue report shows 77 % WER on Catalan.
- Decoding is greedy TDT. Each phrase is decoded in isolation (no context carried over), which is why every phrase gets a capital and a full stop. Cleaning that up (lowercase/strip the period when the next phrase continues the sentence) is an open idea.
3. Segmenting and the CPU cost of listening
- Energy VAD on 30 ms frames. Speech = RMS above
max(250, 3 × noise floor), floor = 10th percentile of the last ~6 s. A phrase ends after 0.7 s of silence; 0.2 s of audio is kept before the first voiced frame and 0.15 s after the last; blips under 0.25 s are dropped. The VAD is crude on purpose: it also fires on keyboard noise and breathing, which decode to empty text and are dropped (harmless, costs a little CPU). - Idle listening cost 13 % of a core at first because the per-frame loop recomputed a percentile for every frame and polled every 100 ms. Fix: compute RMS vectorised, skip the per-frame loop entirely when a chunk is below the threshold and no phrase is open, refresh the floor twice a second, poll at 4 Hz when idle (10 Hz when a phrase is open or dictating). Result: 0.2 %. Decoding is still the real cost, and it runs on any speech in the room.
- Audio older than the open phrase is discarded (
_trim), so capture can run indefinitely at flat memory.
4. Typing: the problem that cost the most time
wtype injects real key events, and Hyprland matches keybinds against the modifier state of all keyboards
combined. If you are still holding SUPER (as you are right after a SUPER+V toggle), typing "Well" fires
SUPER+SHIFT+W, SUPER+E, SUPER+L (lock!)… and SUPER+M would exit the compositor.
Things we tried and what happened:
- Waiting for SUPER to be released, learned from bind press/release events → random 10 s stalls (events arrive from separate client processes in unpredictable order; lost release = wait timeout). Abandoned.
- Reading the keyboard state directly (evdev) → needs the
inputgroup, which is a keylogger-grade permission. Refused. - What works: an empty Hyprland submap. Binds only match in the current submap, so the daemon switches to a
typingsubmap around eachwtypecall and back toresetafterwards (in afinally). The submap needs at least one bind to exist (an emptydefine_submapregisters nothing), so the toggle bind is repeated inside it. The daemon also resets the submap at startup and on SIGTERM so a crash cannot leave your shortcuts dead. Recovery if shortcuts ever stop working:hyprctl dispatch 'hl.dsp.submap("reset")'. - Testing note:
wtype's own modifiers are not combined with bind matching the way a physical SUPER is, so the failure cannot be reproduced withwtype -M logo …. It needs real keys.
Toggle instead of hold-to-talk: we started as push-to-talk, but a hold means SUPER is down the whole time. Also, people release SUPER before V, so a "release of SUPER+V" bind never fires (random lag). A toggle avoids both.
4b. How listening ends (and what must never end it)
Listening is on until one of: the hotkey toggles it off; the explicit spoken phrase "computer, go to sleep" (also "turn off listening" / "disable listening"); the
daemon exits and the session ends (logout, reboot). It never autostarts, so each login begins off, but within a login session the last choice is remembered in $XDG_RUNTIME_DIR/phonon-want-listening, so a daemon restart or crash resumes listening silently (developing/restarting the daemon used to switch it off under the user's feet). It has no idle timeout.
Finishing a command does not end it, and "stop" is silent (no notification). Everyday words ("stop", "mute", "be quiet", "that's all", "thanks")
used to switch listening off entirely (a stray "computer, mute" killed the whole system); they now only end the exchange:
a pending confirm, an armed follow-up after a bare "computer", and any Claude answer still in flight are dropped, and the
daemon goes on waiting for the next wake word. "Mute" mutes the audio. If the capture process dies (PipeWire restart,
input device change, the echo-cancel process going away) a watchdog in Capture._monitor restarts it within ~2 s (rate-limited,
half-read audio dropped); it used to stay "listening" but deaf.
4c. Echo cancellation (music playing while you talk)
The problem is bigger than false wake words. The speech detector's threshold is 3× the noise floor, and steady music
raises the floor. Measured with synthetic music through the laptop speakers: threshold 1782 on the raw mic versus
250 in a quiet room, i.e. speech had to be ~7× louder to register at all, and 27 % of frames counted as "voiced".
Fix: PipeWire's built-in WebRTC echo canceller (integration/phonon-aec.conf, run as its own pipewire -c <conf>
client process; nothing in the system audio config changes, kill it and it is gone). It exposes a virtual mic phonon-mic =
default mic minus what the default speaker is playing (monitor.mode: the speaker's monitor is the reference, so apps keep
playing to the normal default sink). Result: 22 dB less music, threshold back to 250, 5 % of frames over it (transient
residue; blips under 0.25 s are dropped anyway). Noise suppression and gain control are off on purpose (they distort speech).
- Lifecycle: tied to listening. The daemon owns the canceller:
aec_start()when a capture begins (listening or dictation turns on),aec_stop()when the last one ends or the daemon exits (SIGTERM). The canceller costs ~3.9 % of a core and 14 MB while in use, so it is not left running at login. If one is already running (started by hand) the daemon leaves it alone. The pid is kept in$XDG_RUNTIME_DIR/phonon-aec.pid; at startup a canceller left behind by a crashed daemon (SIGKILL test: the orphan survives the crash, then is replaced by one fresh process on restart) is killed first.PHONON_MIC=rawnever starts it. After starting it, the daemon re-checks every 2 s for the virtual mic (first ~1-2 s run on the raw mic). Do not usePR_SET_PDEATHSIGto tie the child to the daemon: it fires when the thread that spawned it exits, and the daemon spawns from short-lived socket-handler threads. - Only while audio is playing. Cancellation costs a little recognition accuracy (the canceller reshapes the mic signal even
when it has nothing to remove), so in
PHONON_MIC=auto(default) the daemon records fromphonon-miconly while the computer is outputting sound and from the raw mic otherwise. "Outputting sound" = any ALSA playback device isRUNNINGin/proc/asound/card*/pcm*p/sub*/status(a few file reads once a second, ~0.3 % CPU; the speaker reads RUNNING while Spotify plays and the HDMI devices stayclosed). It switches to the cancelled mic immediately, and back to the raw mic 8 s after the audio stops (AEC_OFF_DELAY_S; PipeWire also keeps the device open for a few seconds after a stream ends, so the real delay is longer). It never switches mid-phrase, and a switch restarts the capture (~0.2 s, half-read audio dropped). If the status files are missing it assumes audio is playing (keeps cancellation on). Caveats: Bluetooth sinks are not ALSA devices, so for them cancellation stays off (fine for headsets, wrong for a Bluetooth speaker: usePHONON_MIC=aec). Not checked live: that the speaker device really goesclosedwhen Spotify pauses with the canceller attached to its monitor (the logic is tested by simulation; if it stayed RUNNING the result is merely "always cancelled", the old behaviour). - Modes:
PHONON_MIC=auto(above),aec(always the cancelled mic when it exists),raw(never), or a PipeWire node name. Verified: killing the canceller falls back to the raw mic, restarting it switches back. - Not measured: recognition accuracy of real speech over music (the author listens at about half speaking volume; speech over loud music is out of scope), and double-talk behaviour (WebRTC AEC can suppress some of your voice when both are loud).
- Wiring check worth repeating on another machine:
pw-link -l | grep phononmust show the speaker'smonitor_*ports feedingphonon-ec-sinkand the mic feedingphonon-ec-capture; otherwise it cancels nothing and still looks healthy. - Pitfalls hit while testing:
pw-catcannot play a headerless raw file (write a WAV); apgrep -fpattern also matches your own shell (usepgrep -x pw-recordor an anchored pattern like'^pipewire -c').
5. The Claude fallback
- One long-lived
claude -p --input-format stream-json --output-format stream-json --verbose --model haiku --tools ""process. Cold start ≈ 10 s, follow-ups 1–2 s, and it remembers context ("a bit more"). It is pre-warmed when listening is turned on, and recycled after 40 turns / 30 min idle / any error or hang (60 s). - Typed access:
src/phonon-chat.pytalks to the same process through the daemon's socket (chat,chat-run,chat-info,chat-prompt,chat-reset), so there is one session and one shared context, serialised by the session lock (a typed message waits for a spoken one in flight). Every request/reply, voice or typed, is appended toclaude-session.jsonl(git-ignored; it contains what you said after the wake word). Typed/runskips the voice confirm (typing is the confirmation, and there is no mishearing/injection risk) but theBLOCKEDlist still applies; chat runs never trigger auto-promotion. The session cannot be attached to from outside because it uses--no-session-persistence; dropping that flag would make it resumable withclaude --resume, but a resumed TUI session would have tools enabled and would fork the history away from the daemon's live process. - Language bias (Julia). The system prompt tells Claude to use Julia (never Python; there is no numpy here) for scientific
computing, data analysis and every plot:
julia --startup-file=no -e '...'with CairoMakie/DataFrames/Distributions/ Unitful, figures saved under~/Picturesand returned asshow: file. The same preference is stated for Claude Code generally in~/.claude/CLAUDE.md. This needed one safety-check fix: the redirect test inis_risky()ignores quoted strings, otherwise Julia's->lambdas and>comparisons insidejulia -e '...'counted as file redirects and forced a confirmation on nearly every Julia one-liner (echo hi > "$f"is still caught; risky words likerm(still are). - All tools are disabled. It can only propose a shell line as JSON:
{say, shell, show, confirm, repeatable, promote, phrase}. The daemon decides what runs. claude --bareskips login and fails with "Not logged in"; use the plain command.- Safety layers, in order: (1)
BLOCKEDregex refuses plainly destructive/privileged shapes even after confirmation (sudo,rm -rf,mkfs,dd if=, fork bombs,curl|sh, …); (2) Claude says whether confirmation is needed; only an explicitfalseskips it; (3)RISKYregex forces confirmation regardless (rm,mv,kill*,systemctl, package managers,ssh/scp/curl/wget, redirects to files,nmcli … off, …); (4) a spoken "confirm" within 10 s, bare word accepted only while a proposal is pending; "cancel" drops it, and a cancel also invalidates a Claude answer that is still in flight (otherwise it could appear after the cancel). - Prompt-injection / false-trigger reality: a video or another person can say "computer …" and the mic has no echo cancellation. The guards above bound the damage; they do not remove it. Treat listening as opt-in (it is off at every login).
5b. Work mode (elevated session in tmux)
STATUS: working and tested end to end (2026-10-01, see the verification list at the end of this section). Two real bugs
were found only by running it, both fixed: (1) tmux targets: -t =phonon-work is valid for session commands (has-session,
kill-session, set-option, attach) but pane commands (capture-pane, send-keys, paste-buffer) need the trailing colon,
=phonon-work:, else "can't find pane" (screen reads came back empty, so the trust prompt was never answered and nothing could be
sent); (2) an interrupt fires no Stop hook, so the busy flag stuck and the idle timeout could never end the mode.
Why a separate session, and why tmux. Tools are fixed when a Claude process starts, so "elevating" the tool-less command
session is impossible; elevation = a different session with more tools, started explicitly, visible, and time-limited. It runs
the interactive Claude Code UI inside a named tmux session (phonon-work, cwd work/) instead of a hidden headless process, so
you can tmux attach to watch/interrupt/answer prompts, and it is a normal resumable session. (The command session stays
headless: it must return strict JSON in 1-2 s, and driving a screen would be slower and more fragile.)
Mechanics (each verified by hand against Claude Code 2.1.285 and tmux 3.7c):
- Launch:
tmux new-session -d -s phonon-work -c work/ -- claude --model sonnet --permission-mode acceptEdits --tools … --allowed-tools … --disallowed-tools … --append-system-prompt … --settings '<hooks json>'(argv after--, so no shell quoting problems). - First run in a new folder shows "Do you trust this folder?" with "No, exit" preselected: the daemon detects it on screen and sends
Down, Enter. The
SessionStarthook is the readiness signal. - Daemon -> session: single-line text with
tmux send-keys -l, multi-line/long text withload-buffer+paste-buffer -p, then Enter. Quotes,$, backticks, backslashes arrive intact. A multi-line paste is wrapped by the UI in<pasted_content>tags (the model copes); one-liners are not wrapped. A leading/,!,#or@is stripped (slash command / direct shell command that bypasses permissions / memory note). - Session -> daemon: hooks run
phonon-ctl.py work-event <kind>, which writes the hook's JSON payload to a file and sends the daemon<kind> <path>.Stopcarrieslast_assistant_message(no screen scraping) andprompt_id;UserPromptSubmitcarries the prompt;Notificationwithnotification_type: permission_promptmeans a dialog is open;SessionStartcarriessession_id. - Permission dialogs: pressing
1(no Enter) approves that one action, Escape cancels it (and interrupts a running turn). The daemon turns apermission_promptnotification into a pending action, so a spoken "confirm"/"cancel" answers it (60 s). - Permissions measured in a headless session with the same flags: writes inside the cwd allowed, writes outside denied,
juliaallowed,rmdenied, reads outside the cwd denied unlessRead(~/**)is allowed, andRead(~/.ssh/**)etc. stay denied. A caveat that does not go away: the allow/deny lists constrain Claude's Read/Bash tools; Julia is allowed, and Julia code can do anything your account can. No web tools are given, so untrusted web content never shares a session with code execution (research would be a separate, read-only session). - Default-allowed commands: headless sessions auto-approve a built-in list of read-only shell commands (an
echoran without any prompt); do not assume "nothing can run" without a prompt. - A hidden window:
workspace = "special:work silent"in the Hyprland window rule keeps the attached terminal (classphonon-work) on a special workspace without revealing it or taking focus (verified with a throwaway window).SUPER+ALT+Wtoggles it;work-windowreopens it if closed. It closes by itself when the tmux session ends.
State and lifecycle: bar flag 3 = work mode (diamond). It implies listening; turning listening off ends it. Idle limit
(default 3600 s, PHONON_WORK_IDLE_S) counts activity: any hook event (voice, /work, or you typing in the attached window)
resets it, a busy turn never expires. If the daemon restarts while listening resumes and the tmux session still exists, the daemon
adopts it (work mode resumes, idle clock restarts). Killing the tmux session or closing Claude ends the mode within 15 s.
Verified (all against the live daemon): "computer, work mode" by voice -> flag 3, tmux session, hidden special:work window;
first-run trust prompt auto-accepted in a brand-new folder; /work typed round trips ("pong"; a Julia task wrote work/mean.txt
= 50.5); a spoken task's answer delivered by notification; a write outside work/ held, notification quoting the dialog, "confirm"
approved it (file created) and "cancel" declined it (no file); PHONON_WORK_IDLE_S=25 ended the mode on its own; a daemon
restart adopted the running session; "computer stop" interrupts a running turn and clears the busy flag, and a manual Esc in the
window is healed by the watcher in ~26 s; SUPER+ALT+W (work-window) shows/hides the terminal; "computer, normal mode" ends the
session and closes its window; listen-off ends work mode too.
Not verified: the bar diamond's appearance (the state file reads 3; look at it), a task that needs to answer with a file path
(viewer delivery is shared with the command session's tested path), behaviour on newer Claude Code versions (flags, dialog text,
hook payloads were checked against 2.1.285 only), and long unattended runs.
Side effects to know about: answering the trust prompt writes a project entry to ~/.claude.json (the test's temp folder left
one behind); Claude Code runs long shell commands as background shells, so "run X for 90 s" ends the turn at once and the answer
says it is still running; a turn only blocks the idle timeout while it is actually running.
5c. Web research: a separate, web-only agent
Principle: untrusted web content must never share a session with anything that can run code or touch files beyond a report. The voice command session is tool-less (no web, by design); work mode has code execution (no web, by design); research is the third, quarantined piece: web tools and nothing else. The "dual-LLM" idea: the privileged session never reads raw pages, it reads a report written by the quarantined one.
What was measured (Claude Code 2.1.285, Oct 2026):
- A headless session given
--tools WebSearch,WebFetchbut no approval silently acts as if it had no web access ("I don't have web access enabled"). Adding--allowed-tools WebSearch WebFetchmakes it work (it returned Julia v1.13.1 with a source link). - A subagent cannot have tools its parent session lacks: with a parent started
--tools Task,Readand a web-only subagent (--agents), the subagent reported "web tools are unavailable". So "work session without web, subagent with web" is impossible inside one session; hence a separate process. - A simple job takes 20-40 s on Sonnet and yields a structured report: answer first, evidence, sources (with how each was used),
caveats. It reports its own gaps honestly (e.g. "I read the dev-branch NEWS page, not the final notes"). Report file names carry a
made-up time (
…-0000.md/-1200.md): the model does not know the clock; cosmetic.
How it is wired: phonon_research.run() = claude -p --model sonnet --tools WebSearch,WebFetch,Write --allowed-tools WebSearch WebFetch --permission-mode acceptEdits --no-session-persistence --append-system-prompt <rules> with cwd research/ (so Write is
auto-accepted there and nowhere else). Entry points: the phrase table ("research …", "look up …", "search for …"); the command
session's JSON field "research": "<self-contained question>" (set when it needs the web; shell stays empty); the work session's single
allowed Bash command phonon-ctl.py research-work "<question>"; /research in the chat client. Jobs run in daemon threads, at most 2 at
once (research_sem), no spend cap by the user's choice (subscription), only a 30-minute wall-clock guard. Results: deliver_work(text, "Research") (a path on the last line opens in the viewer, the rest is a notification). For a job requested by the work session, the daemon
also types "Research report ready for
Honest limits of the safety story: summarising "in its own words" reduces prompt injection but does not eliminate it; text hidden
in a page can survive into a report. What actually protects you is structural: the research agent cannot execute anything, so a hijacked
one can only write a misleading report; the work session is told to treat reports as data and never to follow instructions in them; work
mode's own limits still apply (permission dialogs, denied curl/rm/ssh); and a hijacked research agent could leak only the question
text through its web requests, so questions must be generic (the work prompt says so) and every request is logged
(research-jobs.jsonl). Cost: each job spends subscription quota; with no cap, a runaway session could use up the window.
Verified end to end: /research typed; a spoken "computer, research …" opening the report in the viewer; the command session
delegating a web question (and not delegating "open the terminal" or arithmetic); the work session, with no web, starting a job through
its allowed command with no permission prompt, receiving the "report ready" message, reading the report and answering correctly.
Not verified: many concurrent jobs, jobs that fail mid-way, very long reports.
6. Learning loop (auto-promotion)
Claude's answer carries repeatable + a generic phrase; repeatable ones are logged to suggested-commands.jsonl
(proposed → ran/failed/cancelled). If it also says promote and the command ran successfully once with no
confirmation needed, auto_promote saves it: a script in commands/<slug>.sh, an entry ("auto": true) in
user-commands.json, a promoted log line, a notification, and a row in AUTO-COMMANDS.md. Guards: never anything that
needed confirmation or matches RISKY/BLOCKED; never a phrase that already exists (built-in, saved, or an app name);
≤ 6 words; ≤ 3 per day (AUTO_PER_DAY). Scripts are plain bash: edit them, or phonon-suggestions.py remove "<phrase>".
Caveat: saved commands keep whatever show mode they had at the time; ones saved before show existed default to
text (our "take a screenshot" had to be re-saved to produce a file).
7. Showing results
show is none | text | file. file: the last stdout line is an absolute path under $HOME or /tmp.
Images → kitten icat in a kitty; PDFs → zathura (needs a PDF plugin; without one you get a notification); text-like
files (≤ 500 KB) → cat in a kitty. text: ≤ 3 lines/240 chars → notification, else a report window. Viewers are kitty
windows of class phonon-output (a Hyprland window rule puts them on special:magic; Hyprland reveals the special
workspace by itself); zathura gets an exec rule and its window address is remembered. SUPER+X kills the viewer
processes the daemon knows about (one process per window).
- Viewers read no input. An early version waited for Enter to close, got focus, and was closed by whatever the user
was typing at the time (the first byte received was an
e, 11 ms after it opened). A no-focus version was tried and then dropped at the author's request; normal focus +SUPER+Qis the current behaviour. - A viewer only ever opens files under
$HOMEor/tmp.
8. Running commands: run()
grim - | wl-copyleaveswl-copyalive in the background holding the clipboard; it inherits the output pipe.subprocess.run(capture_output=True)then waits for EOF until its timeout and reports failure for a command that worked. Output now goes to a temp file and only the process is waited on.- A timeout that kills is wrong here: app launches (
kitty, a browser) run until you close them and would be killed. So: app launches are detached (never waited on), everything else waits up to 60 s and then just reports "still running". Nothing is killed for being slow.
9. Gotchas (short list)
| Symptom | Cause / fix |
|---|---|
| All shortcuts dead | Stuck in the typing submap: hyprctl dispatch 'hl.dsp.submap("reset")'. |
| Listening indicator on but nothing happens | Should self-heal now (capture watchdog, see capture restarted in the log). If not: phonon-ctl.py listen-off then listen-on. |
| Mic seems live | Bar dot/ring is the only sign; src/phonon-ctl.py listen / toggle end it. Flag file phonon-recording. |
pkill -f phonon… kills your own shell | The pattern matches the shell's own command line. Kill by PID: pgrep -f "^…/venv/bin/python". |
| Command works but reports failure after 60 s | Pipe held open by a background child (§8). |
| Saved command opens no window | Saved before show existed; set "show": "file" in user-commands.json. |
| Phrase typed a few seconds late, in pairs | (old bug) SUPER tracking, see §4. |
New .qml file not picked up | Quickshell reload race; restart Quickshell. Editing existing files hot-reloads. |
| Weights fail to load | Check vendor/fermion_container.py matches the upstream checksum (setup.sh does). |
10. Testing
There is no automated test suite. Useful manual hooks: phonon-ctl.py say "computer go to workspace two" pretends that
phrase was heard (exercises everything after the microphone); phonon-ctl.py show <file> opens a file in the viewer;
phonon_commands.py is pure text-in/Action-out and can be imported and tested with no model:
python3 -c "import phonon_commands as c; print(c.match('go to workspace two'))" (run from src/).
Watch $XDG_RUNTIME_DIR/phonon.log (per-phrase decode times, cuts, commands; ignored phrases are logged by length only).
11. Privacy
The mic is open only while dictating or listening. Ignored phrases are never logged as text. Dictated text is written to
$XDG_RUNTIME_DIR/phonon-last.txt and the log (tmpfs, cleared at reboot). Only the text after the wake word can reach
Claude. suggested-commands.jsonl persists the commands you spoke and what they became (git-ignored).
12. Open ideas
Rolling partial results with backspace correction (more "live", fragile in arbitrary apps) · joining phrases into
sentences (capitalisation/period) · echo cancellation or media-aware ducking for the wake word · a real wake-word model
instead of "transcribe everything" · HTML reports rendered instead of shown as source · a PDF fallback via pdftoppm +
icat while no zathura plugin exists · porting the Hyprland strings behind a small adapter · automated tests.