Skip to content

Latest commit

 

History

History
57 lines (47 loc) · 17.2 KB

File metadata and controls

57 lines (47 loc) · 17.2 KB

AGENTS.md

Python CLI tool for local speech-to-text using whisper.cpp in Docker. Stdlib only, no pip dependencies. Package: digue/.

Commands

  • Test: pytest tests/ -v --tb=short (or make test)
  • Type check: mypy (strict, configured in pyproject; or make mypy)
  • Lint: ruff check . --fix && ruff format --line-length 120 (or make lint)
  • Run: python -m digue dictate (toggle dictation), python -m digue detect etc. Console script entry: digue.cli:main.
  • Smoke test after building a wheel: make smoke-wheel. It installs the wheel into a temporary virtual environment, with no runtime dependencies, and runs the installed digue --version and digue config show entrypoint.
  • All of the above have make targets (make help); make check runs lint-check + mypy + test.

Conventions

  • Stdlib only. No external runtime dependencies. tomllib (3.11+), urllib.request, subprocess, pathlib.
  • Package. Submodules live under digue/. Version is digue.__version__. Run via python -m digue or the digue console script (digue.cli:main). Public import digue surface (__all__): __version__, load_config, transcribe_file, record_to. Modules: config, notify, container, transcribe, convert, recording, delivery, audio, dictate, benchmark, cli, __main__.
  • English everywhere. README, docstrings, comments, CLI help text, notifications, commit messages -- all English.
  • pathlib.Path always. Never os.path.
  • Modern type hints. str | None, list[Path] -- not Optional, List.
  • No single-char variables except _ in unpacking.
  • Lazy imports. Each package module (except __main__.py) may import at module level only argparse, collections.abc, contextlib, dataclasses, os, pathlib, sys, typing, and digue (sibling modules / package constants). Everything else is imported inside the function that uses it, and a function never re-imports a module-level name (os, contextlib, Path were re-imported 42 times before the test caught it). The test suite scans every digue/*.py except __main__.py.
  • Type hints everywhere; mypy --strict must stay clean. Config dicts are dict[str, dict[str, Any]] (values are TOML-parsed scalars; Any beats object because dict is invariant and these values flow into str/int params).
  • Errors are visible. send_notification() always prints to stderr AND tries desktop notification. Never silently swallow errors.
  • User-facing failures notify, never traceback. A failure inside a dictation/transcription flow (save, transcribe, paste) reports via send_notification(..., timeout_ms=...) and returns exit code 1; raw tracebacks are for bugs only.

Architecture decisions

  • Native server formats: wav, flac, mp3, ogg/Vorbis, aiff (verified empirically against whisper-server built with WHISPER_COMMON_FFMPEG=OFF; opus-in-ogg, m4a/AAC, mp4, webm, mka, wma fail with HTTP 400). Unsupported formats are converted in memory (ffmpeg writes to stdout pipe, nothing hits the disk), upfront for known-bad extensions and as a retry after HTTP 400.
  • Every request sends token_timestamps=false: without it, the server enables token timestamps for text format, which triggers whisper's max_len=60 segment wrapping on token boundaries -- this split words in half across lines ("trans"/"crevendo"). Verified: with the flag, output comes as natural segments.
  • backend = "remote" means the server runs on another machine: via SSH tunnel (default, remote-host empty = 127.0.0.1) or directly on the LAN (remote-host set to a host/IP; port stays server.port). digue must never create, start, or stop a local container for it; all container management commands refuse. When the remote host is the default (tunnel), errors suggest the ssh command; with a LAN host, they show host:port instead.
  • Dictation runs as a daemon: the first digue dictate starts the recorder and stays alive waiting (200ms poll); the second toggle signals SIGTERM and exits instantly (the daemon stops the recorder and delivers: transcribe -> paste -> txt+flac). In a terminal, Ctrl+c does the same (the daemon installs its own SIGINT handler -- the global KeyboardInterrupt handler would discard the take and leave the recorder running), and stderr announces "Press Ctrl+c to stop recording and transcribe". The daemon enforces the duration limit itself and replaces the "Recording..." popup with "Limit reached (s), transcribing..." by id (per-take id = base + pid % 32, so overlapping takes never replace or close each other's popups). The detached watchdog is only a safety killer (kill, no notify, no re-run) for a SIGKILLed daemon, and it sleeps WATCHDOG_GRACE_SECONDS past the limit: with an equal deadline it won the race against the 200ms poll (measured at 20s) and the daemon saw "died" instead of "limit"; the daemon also reports any recorder exit at or past the limit as "limit". After the recorder stops, the SIGINT/SIGTERM handlers stay installed on purpose: a second Ctrl+c or a pkill digue during delivery is ignored (the flags are only read while recording), so a take is never dropped mid-delivery; only SIGKILL aborts it (documented in the README, Recording). Overlapping takes are supported: while a daemon delivers (daemon pid file state "delivering"), a new toggle starts a fresh take; each daemon stops only its own recorder (stop_recording_pid, capture the rec_file while the recorder is alive). A recorder alive with no daemon at all (pkill digue) is recovered by the next toggle (stop + deliver what kept recording). The daemon file stores <pid> <state> <starttime>: a pid is only an identity together with its /proc starttime (the file outlives a SIGKILLed daemon and pids get recycled), so a toggle never signals a pid whose starttime differs; a file without starttime is treated as absent. State files (daemon state, take state) are published with _write_state_file (temp sibling + rename): Path.write_text truncates first, and a toggle reading that empty window would see "no daemon" and start a second take. There is no global recording state: _pid_file/is_recording/stop_recording/_recorder_pid_file and the newest-WAV fallback were removed, and the take state JSON is the only per-take identity, so a stop can only ever act on the recorder of the take being handled. pw-record has no --duration flag (checked man page); arecord has no limit at all.
  • Saved dictation audio is compressed with audio-format (default flac): flac is lossless and natively decodable by whisper-server, so retranscription never needs the ffmpeg fallback; opus is ~7% of WAV but lossy and forces the fallback. If compression fails (or ffmpeg is missing), keep the WAV and warn -- never leave a partial compressed file in its place. Saved names carry the take id (<YYYYMMDD-HHMMSS>-<take_id>.<ext>) so two takes ending in the same second never overwrite each other, and every save is exclusive ("xb"/reservation up front + temp file + replace()): a collision raises and the WAV is rescued instead of silently overwritten. Delivery comes first: transcription and paste/type run on the live WAV, and the archiving (copy + ffmpeg compression, the slow part) runs only after the text was delivered; if anything fails after the recording stopped, the raw recording is rescued (moved) to <audio_dir>/YYYY/MM/<timestamp>-<take_id>.<ext> (.wav, or .flac for a native FLAC take) so no take is lost -- save-audio = false does not change this: it only skips the backup of a delivered take.
  • start_recording publishes the take state (digue-take-<take_id>.json) before starting the recorder: starting (reserved rec_file + take id) -> Popen (failure removes the state: no WAV, no process) -> recording (recorder pid + /proc starttime) -> only then the watchdog spawn. Identity comes before the watchdog on purpose: publishing the state is two syscalls (~50 us) while spawning the watchdog is fork+exec (~10 ms), so the recorder-without-identity window shrinks to ~50 us and the no-watchdog window to ~10 ms -- and a recorder with identity is the case recovery can stop on its own. Residual windows, documented and accepted (an intermediate launcher would remove both and is not worth the cost): a daemon killed between Popen and publishing recording (~50 us) leaves a recorder with no identity and no watchdog -- only the user can stop it (pkill pw-record); a daemon killed between publishing recording and spawning the watchdog (~10 ms) leaves a recorder without a duration limit but with identity, and the next toggle stops and delivers it. In recovery an empty WAV is terminal whatever the take's age (stop_recording_pid unlinks it and the recorder is dead by then; the no-speech path archives it): a young-empty guard once kept the state of a take with no audio, so every toggle for a minute reclaimed it, reported "Empty or missing audio file" and blocked the surplus rescue -- and its tests missed it because they mocked finish_dictation whole. Tests of the take-state lifecycle must run the real finish_dictation (mock transcribe, send_text, send_notification, save_audio; they isolate everything external) so they see the WAV disappear. An orphaned starting state (daemon died between publishing and registering the recorder) is handled conservatively by the next toggle, with no /proc/*/fd scanning: a state younger than ORPHAN_MIN_AGE_SECONDS (60 s) is not even claimed (a claim always means work; the toggle records normally); after that a missing/empty WAV expires together with its state, and a non-empty WAV is rescued (never transcribed/pasted automatically -- the recorder may still be writing) with a warning. The daemon removes its take state only after finish_dictation returns a terminal outcome (the daemon file always goes: the process is exiting); a retryable failure or an unexpected exception keeps state and WAV, and the next toggle claims the take. Every post-paste outcome is terminal (a .txt or archive failure after a successful paste is delivered/rescued with exit 1), so a retry never pastes twice. Recovery of a take in delivering could otherwise paste twice (daemon died between the paste and the state removal, i.e. during the archive): finish_dictation writes the .txt right after pasting, and _recover_claimed_take looks for *-<take_id>.txt in audio-dir first -- when it exists the text was delivered, and only the audio is archived (_archive_recovered_take, next to the .txt). Dying during the archive usually leaves a partial .wav copy, an empty compressed reservation or a .tmp with that same stem next to the .txt; _archive_recovered_take removes them first (same stem = same take id, and the live WAV is complete), otherwise the exclusive copy and the rescue both collide and the good audio stays in the runtime dir with no state. The residual window is the few ms between send_text and the .txt write. The take state is the lifecycle source of truth: starting -> recording -> delivering while the owner lives (_mark_take_delivering runs with the daemon file transition, before the recorder is stopped; the daemon file tracks the same phases, and during delivery a second toggle must not stop anything and starts a new take instead); when the owner dies, the next toggle claims the orphan under the dictate lock, rewrites it to recovering with its own pid+starttime, delivers it and returns (it never records: it publishes no daemon state, so a concurrent toggle starts a new take; recording after the recovery would leave two recorders competing for one daemon state, with a single stop), removes it after a terminal outcome (delivered, rescued, empty) and keeps it through retryable_failure or an unexpected exception, so the next toggle retries. A malformed take state JSON is reported as "unreadable" on stderr and never removed (the WAV stays).
  • After the oldest orphan take is delivered, the remaining orphans are claimed one by one (each claim under the dictate lock) and rescued: a live recorder is stopped through its published identity, the recording is moved to <audio_dir>/YYYY/MM/<timestamp>-<take_id>.<ext> (live suffix kept) and the take state JSON is archived next to it as <timestamp>-<take_id>.json with state = "rescued" (metadata of the recording); one consolidated "N recordings rescued to " notification covers all of them. Only the oldest orphan is ever transcribed and pasted. A failed rescue keeps the take state for the next toggle. clean treats the .json as metadata of the recording, not a category: it is removed together with the recording of the same stem (counted as one unit), never listed as a transcript, and a .json whose recording is gone is preserved.
  • Default data dir is $XDG_DATA_HOME/digue (~/.local/share/digue); XDG_DATA_HOME is often unset, so the fallback to ~/.local/share matters -- do not assume it is set.
  • Config is split by role: [transcribe] is shared by transcribe, batch-transcribe and dictate for every applicable transcription-time option (language, prompt, output-format, cue wrapping); CLI options override it. [dictate] holds capture/delivery options (audio-dir, recorder, input-mode, ...). No option lives in the wrong section.
  • Per-machine config lives in [host.<hostname>][section] tables (defaults < global < host, hostname matched with or without its domain part); one config.toml can be versioned in dotfiles for all machines. gethostname() is in-memory (~4us), safe to call on every run.
  • Docker images have CPU instruction compatibility issues. main crashes on Meteor Lake (AMX), main-vulkan crashes on Kaby Lake (SIGILL). The config allows overriding image per machine, and digue server start --image for one run. docker start reuses the container's original image, so cmd_server_start compares it (container_image, docker inspect {{.Config.Image}}) with the resolved one and recreates on mismatch; ensure_server (hotkey path) only warns -- a multi-GB pull is not what a keypress asked for. The Docker container is named digue-whisper.cpp by default (server.container-name, digue server start -n). See DOCKER_IMAGES dict and comments.
  • Notifications have a lifecycle contract: each process uses slot NOTIFY_REPLACE_ID + pid % NOTIFY_ID_SLOTS, and every progress notification must eventually be replaced or closed without affecting overlapping takes. Successful dictation replaces the progress popup with a 3s "Pasted (N chars)" / "Typed (N chars)" toast (path still goes to stderr); main() calls notify_close() on KeyboardInterrupt. Failures replace progress with a timed error (5-10s). Only a kill -9 can leave one stuck; the README manual fix closes all slots.
  • send_notification() uses --replace-id with the process slot so each take replaces only its own notification. Default timeout is 0 (stays until replaced). Only success/error messages get timeouts.
  • xclip must be called with stdout=DEVNULL, stderr=DEVNULL (not capture_output=True) because it forks a background process that inherits pipes and causes timeout.
  • Terminal progress output must be TTY-aware: _stderr_is_tty() gates everything. On a TTY, progress redraws one line with \r (bar + notification share the line via redraw). On captured/piped stderr, \r does nothing visually -- each print would become a huge line in logs -- so progress prints sparse plain lines (one per ~5%) with no bar, and send_notification() prints plain lines with \n. Never assume stderr is a TTY; every progress/notification print must handle both paths.
  • create_container auto-downloads the model if missing, to avoid Docker crash loops.
  • AVAILABLE_MODELS is the closed list of every ggml-*.bin in huggingface.co/ggerganov/whisper.cpp (f16, -q8_0, -q5_0/-q5_1, .en), and benchmark.MODEL_SIZES_MB must list the same names in the same order (a test checks both). The name is the download URL and the container's --model, so a typo would create a crash-looping container: never accept free-form names. Quantized models mainly save disk/RAM; speed gains are CPU-side and must be benchmarked, not assumed.
  • doctor never pulls Docker images; compatibility tests run only for images already present locally.
  • Publishing: version has a single source of truth, digue.__version__ (bump it before building). make build clears dist/ first; before uploading, run make build-check, which also smoke-tests the installed wheel (see Commands).

Testing

  • Tests use pytest with unittest.mock for subprocess/network calls. No real Docker or network in tests.
  • Test classes group related tests (class TestDetectBackend:, class TestNotify:).
  • All create_container tests must mock download_model and pull_image too, or they'll attempt real network downloads or Docker pulls.
  • Tests that mock save_audio must set return_value=("path", "timestamp") -- callers unpack the tuple.
  • Lifecycle tests for take state must not mock finish_dictation whole: its internal steps (recorder stop unlinks an empty WAV, no-speech path archives the audio) are exactly what the assertions are about -- Bug 1 of the 2026-09-05 review slipped through because two tests mocked it and never saw the WAV disappear. Mocking transcribe/send_text/send_notification isolates what tests need while running the real delivery flow.

If you find a wrong assumption in this file during a session, suggest the correction.