Python CLI tool for local speech-to-text using whisper.cpp in Docker. Stdlib only, no pip dependencies. Package:
digue/.
- Test:
pytest tests/ -v --tb=short(ormake test) - Type check:
mypy(strict, configured in pyproject; ormake mypy) - Lint:
ruff check . --fix && ruff format --line-length 120(ormake lint) - Run:
python -m digue dictate(toggle dictation),python -m digue detectetc. Console script entry:digue.cli:main. - Smoke test after building a wheel:
make smoke-wheel. It installs the wheel into a temporary virtual environment, with no runtime dependencies, and runs the installeddigue --versionanddigue config showentrypoint. - All of the above have
maketargets (make help);make checkruns lint-check + mypy + test.
- Stdlib only. No external runtime dependencies.
tomllib(3.11+),urllib.request,subprocess,pathlib. - Package. Submodules live under
digue/. Version isdigue.__version__. Run viapython -m digueor thedigueconsole script (digue.cli:main). Publicimport diguesurface (__all__):__version__,load_config,transcribe_file,record_to. Modules:config,notify,container,transcribe,convert,recording,delivery,audio,dictate,benchmark,cli,__main__. - English everywhere. README, docstrings, comments, CLI help text, notifications, commit messages -- all English.
pathlib.Pathalways. Neveros.path.- Modern type hints.
str | None,list[Path]-- notOptional,List. - No single-char variables except
_in unpacking. - Lazy imports. Each package module (except
__main__.py) may import at module level onlyargparse,collections.abc,contextlib,dataclasses,os,pathlib,sys,typing, anddigue(sibling modules / package constants). Everything else is imported inside the function that uses it, and a function never re-imports a module-level name (os,contextlib,Pathwere re-imported 42 times before the test caught it). The test suite scans everydigue/*.pyexcept__main__.py. - Type hints everywhere;
mypy --strictmust stay clean. Config dicts aredict[str, dict[str, Any]](values are TOML-parsed scalars;Anybeatsobjectbecausedictis invariant and these values flow intostr/intparams). - Errors are visible.
send_notification()always prints to stderr AND tries desktop notification. Never silently swallow errors. - User-facing failures notify, never traceback. A failure inside a dictation/transcription flow (save, transcribe, paste) reports via
send_notification(..., timeout_ms=...)and returns exit code 1; raw tracebacks are for bugs only.
- Native server formats: wav, flac, mp3, ogg/Vorbis, aiff (verified empirically against whisper-server built with WHISPER_COMMON_FFMPEG=OFF; opus-in-ogg, m4a/AAC, mp4, webm, mka, wma fail with HTTP 400). Unsupported formats are converted in memory (ffmpeg writes to stdout pipe, nothing hits the disk), upfront for known-bad extensions and as a retry after HTTP 400.
- Every request sends
token_timestamps=false: without it, the server enables token timestamps for text format, which triggers whisper's max_len=60 segment wrapping on token boundaries -- this split words in half across lines ("trans"/"crevendo"). Verified: with the flag, output comes as natural segments. backend = "remote"means the server runs on another machine: via SSH tunnel (default,remote-hostempty = 127.0.0.1) or directly on the LAN (remote-hostset to a host/IP;portstaysserver.port).diguemust never create, start, or stop a local container for it; all container management commands refuse. When the remote host is the default (tunnel), errors suggest the ssh command; with a LAN host, they show host:port instead.- Dictation runs as a daemon: the first
digue dictatestarts the recorder and stays alive waiting (200ms poll); the second toggle signals SIGTERM and exits instantly (the daemon stops the recorder and delivers: transcribe -> paste -> txt+flac). In a terminal, Ctrl+c does the same (the daemon installs its own SIGINT handler -- the global KeyboardInterrupt handler would discard the take and leave the recorder running), and stderr announces "Press Ctrl+c to stop recording and transcribe". The daemon enforces the duration limit itself and replaces the "Recording..." popup with "Limit reached (s), transcribing..." by id (per-take id = base + pid % 32, so overlapping takes never replace or close each other's popups). The detached watchdog is only a safety killer (kill, no notify, no re-run) for a SIGKILLed daemon, and it sleepsWATCHDOG_GRACE_SECONDSpast the limit: with an equal deadline it won the race against the 200ms poll (measured at 20s) and the daemon saw "died" instead of "limit"; the daemon also reports any recorder exit at or past the limit as "limit". After the recorder stops, the SIGINT/SIGTERM handlers stay installed on purpose: a second Ctrl+c or apkill digueduring delivery is ignored (the flags are only read while recording), so a take is never dropped mid-delivery; only SIGKILL aborts it (documented in the README, Recording). Overlapping takes are supported: while a daemon delivers (daemon pid file state "delivering"), a new toggle starts a fresh take; each daemon stops only its own recorder (stop_recording_pid, capture the rec_file while the recorder is alive). A recorder alive with no daemon at all (pkill digue) is recovered by the next toggle (stop + deliver what kept recording). The daemon file stores<pid> <state> <starttime>: a pid is only an identity together with its /proc starttime (the file outlives a SIGKILLed daemon and pids get recycled), so a toggle never signals a pid whose starttime differs; a file without starttime is treated as absent. State files (daemon state, take state) are published with_write_state_file(temp sibling + rename):Path.write_texttruncates first, and a toggle reading that empty window would see "no daemon" and start a second take. There is no global recording state:_pid_file/is_recording/stop_recording/_recorder_pid_fileand the newest-WAV fallback were removed, and the take state JSON is the only per-take identity, so a stop can only ever act on the recorder of the take being handled. pw-record has no--durationflag (checked man page); arecord has no limit at all. - Saved dictation audio is compressed with
audio-format(defaultflac): flac is lossless and natively decodable by whisper-server, so retranscription never needs the ffmpeg fallback; opus is ~7% of WAV but lossy and forces the fallback. If compression fails (or ffmpeg is missing), keep the WAV and warn -- never leave a partial compressed file in its place. Saved names carry the take id (<YYYYMMDD-HHMMSS>-<take_id>.<ext>) so two takes ending in the same second never overwrite each other, and every save is exclusive ("xb"/reservation up front + temp file +replace()): a collision raises and the WAV is rescued instead of silently overwritten. Delivery comes first: transcription and paste/type run on the live WAV, and the archiving (copy + ffmpeg compression, the slow part) runs only after the text was delivered; if anything fails after the recording stopped, the raw recording is rescued (moved) to<audio_dir>/YYYY/MM/<timestamp>-<take_id>.<ext>(.wav, or.flacfor a native FLAC take) so no take is lost --save-audio = falsedoes not change this: it only skips the backup of a delivered take. start_recordingpublishes the take state (digue-take-<take_id>.json) before starting the recorder:starting(reserved rec_file + take id) ->Popen(failure removes the state: no WAV, no process) ->recording(recorder pid + /proc starttime) -> only then the watchdog spawn. Identity comes before the watchdog on purpose: publishing the state is two syscalls (~50 us) while spawning the watchdog is fork+exec (~10 ms), so the recorder-without-identity window shrinks to ~50 us and the no-watchdog window to ~10 ms -- and a recorder with identity is the case recovery can stop on its own. Residual windows, documented and accepted (an intermediate launcher would remove both and is not worth the cost): a daemon killed betweenPopenand publishingrecording(~50 us) leaves a recorder with no identity and no watchdog -- only the user can stop it (pkill pw-record); a daemon killed between publishingrecordingand spawning the watchdog (~10 ms) leaves a recorder without a duration limit but with identity, and the next toggle stops and delivers it. In recovery an empty WAV is terminal whatever the take's age (stop_recording_pidunlinks it and the recorder is dead by then; the no-speech path archives it): a young-empty guard once kept the state of a take with no audio, so every toggle for a minute reclaimed it, reported "Empty or missing audio file" and blocked the surplus rescue -- and its tests missed it because they mockedfinish_dictationwhole. Tests of the take-state lifecycle must run the realfinish_dictation(mocktranscribe,send_text,send_notification,save_audio; they isolate everything external) so they see the WAV disappear. An orphanedstartingstate (daemon died between publishing and registering the recorder) is handled conservatively by the next toggle, with no/proc/*/fdscanning: a state younger thanORPHAN_MIN_AGE_SECONDS(60 s) is not even claimed (a claim always means work; the toggle records normally); after that a missing/empty WAV expires together with its state, and a non-empty WAV is rescued (never transcribed/pasted automatically -- the recorder may still be writing) with a warning. The daemon removes its take state only afterfinish_dictationreturns a terminal outcome (the daemon file always goes: the process is exiting); a retryable failure or an unexpected exception keeps state and WAV, and the next toggle claims the take. Every post-paste outcome is terminal (a.txtor archive failure after a successful paste isdelivered/rescuedwith exit 1), so a retry never pastes twice. Recovery of a take indeliveringcould otherwise paste twice (daemon died between the paste and the state removal, i.e. during the archive):finish_dictationwrites the.txtright after pasting, and_recover_claimed_takelooks for*-<take_id>.txtin audio-dir first -- when it exists the text was delivered, and only the audio is archived (_archive_recovered_take, next to the.txt). Dying during the archive usually leaves a partial.wavcopy, an empty compressed reservation or a.tmpwith that same stem next to the.txt;_archive_recovered_takeremoves them first (same stem = same take id, and the live WAV is complete), otherwise the exclusive copy and the rescue both collide and the good audio stays in the runtime dir with no state. The residual window is the few ms betweensend_textand the.txtwrite. The take state is the lifecycle source of truth:starting -> recording -> deliveringwhile the owner lives (_mark_take_deliveringruns with the daemon file transition, before the recorder is stopped; the daemon file tracks the same phases, and during delivery a second toggle must not stop anything and starts a new take instead); when the owner dies, the next toggle claims the orphan under the dictate lock, rewrites it torecoveringwith its own pid+starttime, delivers it and returns (it never records: it publishes no daemon state, so a concurrent toggle starts a new take; recording after the recovery would leave two recorders competing for one daemon state, with a single stop), removes it after a terminal outcome (delivered,rescued,empty) and keeps it throughretryable_failureor an unexpected exception, so the next toggle retries. A malformed take state JSON is reported as "unreadable" on stderr and never removed (the WAV stays).- After the oldest orphan take is delivered, the remaining orphans are claimed one by one (each claim under the dictate lock) and rescued: a live recorder is stopped through its published identity, the recording is moved to
<audio_dir>/YYYY/MM/<timestamp>-<take_id>.<ext>(live suffix kept) and the take state JSON is archived next to it as<timestamp>-<take_id>.jsonwithstate = "rescued"(metadata of the recording); one consolidated "N recordings rescued to " notification covers all of them. Only the oldest orphan is ever transcribed and pasted. A failed rescue keeps the take state for the next toggle.cleantreats the .json as metadata of the recording, not a category: it is removed together with the recording of the same stem (counted as one unit), never listed as a transcript, and a .json whose recording is gone is preserved. - Default data dir is
$XDG_DATA_HOME/digue(~/.local/share/digue); XDG_DATA_HOME is often unset, so the fallback to~/.local/sharematters -- do not assume it is set. - Config is split by role:
[transcribe]is shared by transcribe, batch-transcribe and dictate for every applicable transcription-time option (language, prompt, output-format, cue wrapping); CLI options override it.[dictate]holds capture/delivery options (audio-dir, recorder, input-mode, ...). No option lives in the wrong section. - Per-machine config lives in
[host.<hostname>][section]tables (defaults < global < host, hostname matched with or without its domain part); one config.toml can be versioned in dotfiles for all machines.gethostname()is in-memory (~4us), safe to call on every run. - Docker images have CPU instruction compatibility issues.
maincrashes on Meteor Lake (AMX),main-vulkancrashes on Kaby Lake (SIGILL). The config allows overridingimageper machine, anddigue server start --imagefor one run.docker startreuses the container's original image, socmd_server_startcompares it (container_image,docker inspect {{.Config.Image}}) with the resolved one and recreates on mismatch;ensure_server(hotkey path) only warns -- a multi-GB pull is not what a keypress asked for. The Docker container is nameddigue-whisper.cppby default (server.container-name,digue server start -n). SeeDOCKER_IMAGESdict and comments. - Notifications have a lifecycle contract: each process uses slot
NOTIFY_REPLACE_ID + pid % NOTIFY_ID_SLOTS, and every progress notification must eventually be replaced or closed without affecting overlapping takes. Successful dictation replaces the progress popup with a 3s "Pasted (N chars)" / "Typed (N chars)" toast (path still goes to stderr);main()callsnotify_close()onKeyboardInterrupt. Failures replace progress with a timed error (5-10s). Only akill -9can leave one stuck; the README manual fix closes all slots. send_notification()uses--replace-idwith the process slot so each take replaces only its own notification. Default timeout is 0 (stays until replaced). Only success/error messages get timeouts.xclipmust be called withstdout=DEVNULL, stderr=DEVNULL(notcapture_output=True) because it forks a background process that inherits pipes and causes timeout.- Terminal progress output must be TTY-aware:
_stderr_is_tty()gates everything. On a TTY, progress redraws one line with\r(bar + notification share the line via redraw). On captured/piped stderr,\rdoes nothing visually -- each print would become a huge line in logs -- so progress prints sparse plain lines (one per ~5%) with no bar, andsend_notification()prints plain lines with\n. Never assume stderr is a TTY; every progress/notification print must handle both paths. create_containerauto-downloads the model if missing, to avoid Docker crash loops.AVAILABLE_MODELSis the closed list of everyggml-*.binin huggingface.co/ggerganov/whisper.cpp (f16,-q8_0,-q5_0/-q5_1,.en), andbenchmark.MODEL_SIZES_MBmust list the same names in the same order (a test checks both). The name is the download URL and the container's--model, so a typo would create a crash-looping container: never accept free-form names. Quantized models mainly save disk/RAM; speed gains are CPU-side and must be benchmarked, not assumed.doctornever pulls Docker images; compatibility tests run only for images already present locally.- Publishing: version has a single source of truth,
digue.__version__(bump it before building).make buildclearsdist/first; before uploading, runmake build-check, which also smoke-tests the installed wheel (see Commands).
- Tests use
pytestwithunittest.mockfor subprocess/network calls. No real Docker or network in tests. - Test classes group related tests (
class TestDetectBackend:,class TestNotify:). - All
create_containertests must mockdownload_modelandpull_imagetoo, or they'll attempt real network downloads or Docker pulls. - Tests that mock
save_audiomust setreturn_value=("path", "timestamp")-- callers unpack the tuple. - Lifecycle tests for take state must not mock
finish_dictationwhole: its internal steps (recorder stop unlinks an empty WAV, no-speech path archives the audio) are exactly what the assertions are about -- Bug 1 of the 2026-09-05 review slipped through because two tests mocked it and never saw the WAV disappear. Mockingtranscribe/send_text/send_notificationisolates what tests need while running the real delivery flow.
If you find a wrong assumption in this file during a session, suggest the correction.