Files
max-index/docs/goals/max-index-v1/PLAN.md

233 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Max Index V1 Implementation Plan
**Intent:** превратить пустой репозиторий в воспроизводимое кроссплатформенное desktop-приложение, которое локально распознаёт русскую речь и один раз увеличивает индекс при подтверждённой фразе.
**Current Behavior:** Phase 0 уже даёт packaged Electron/React shell, закреплённый dependency graph и чистые index/face primitives; захват микрофона, recognizer, оркестрация сессии и редактируемые правила ещё не реализованы.
**Expected Outcome:** разработчик поднимает проект через `pnpm`, а пользователь получает честно обозначенное локальное прослушивание, реакцию персонажа и индекс `0..100` без записи сырого звука и полной стенограммы.
**Target-Perspective Output:** пользователь запускает приложение, выбирает микрофон, нажимает Start, произносит настроенную фразу и видит ровно одно изменение индекса и короткую реакцию; Pause действительно освобождает микрофон, повторный запуск после загрузки модели работает offline.
**Truth Owner:** требования принадлежат корневому `AGENTS.md`; зависимости — `package.json` и `pnpm-lock.yaml`; желаемое состояние сессии, индекс и история — экземпляр `SessionController` в main-процессе поверх чистого domain reducer; MediaStream — renderer; готовность recognizer, очередь и границы utterance — utility process.
**Contract Boundary:** узкий типизированный preload API для управления и versioned snapshots; `RecognitionEvent` с `sessionId`, `utteranceId`, `sequence`, `kind` и `text`; отдельный bounded `MessagePort` для будущего PCM data plane.
**Cutover:** greenfield; mock recognizer сначала доказывает настоящую межпроцессную границу, затем `sherpa-onnx` становится единственным production recognizer.
**Displaced Path:** удалить любой Forge demo-код; после подключения sherpa оставить mock только как test/dev dependency injection; после появления оригинальных ассетов удалить CSS-заглушку лица.
**Value Density:** первый законченный срез проходит через renderer, preload, main, utility и domain, доказывая главную инварианту partial → final без двойного увеличения.
**Acceptance Evidence:** на целевой платформе виден сценарий «микрофон → русская фраза → одно увеличение → Pause останавливает tracks → offline restart»; unit/integration/smoke проверки покрывают границы процессов и отказ recognizer.
**Evidence Lane:** детерминированные Vitest-тесты, packaged smoke на текущем хосте, ручной сценарий с реальным микрофоном и отдельная release-матрица по OS/arch.
**Kill Criteria:** packaged build не содержит доступного пользователю mock-пути; PCM не проходит через control IPC; renderer не хранит канонический индекс; сырые аудиофайлы и полный transcript не появляются на диске; system loopback не смешан с microphone provider.
**Architecture Slice:** `Renderer capture → MessagePort → Utility recognizer → RecognitionEvent → Main SessionController/domain → versioned snapshot → Renderer projection`.
**Plan Review Gate:** PRE passed for Phase 0 on 2026-07-10; later phases require task-level PRE review before execution.
## Границы результата
В V1 входят микрофон, локальный streaming ASR, правила, индекс, история текущей сессии, лицо, Start/Pause/Reset, выбор устройства, tray и установщики. Аккаунты, backend, полная стенограмма, запись встреч, speaker identification и system loopback остаются вне V1.
Текущий запрос реализует только Phase 0. Он создаёт запускаемый фундамент и не выдаёт статическую оболочку за работающий recognizer.
## Карта архитектуры
### Source of truth
- `AGENTS.md` — продуктовые и инженерные ограничения.
- `package.json` + `pnpm-lock.yaml` — воспроизводимый dependency graph.
- `SessionController` в main — desired state, `sessionId`, versioned snapshot, индекс и короткая история текущей сессии.
- Pure domain reducer — matching, dedupe, cooldown, clamp и выбор face level.
- Renderer — только MediaStream handle/capture health и отображение snapshot.
- Utility process — recognizer readiness, endpointing, `utteranceId`, sequence и bounded PCM queue.
- Main model manager — версия модели, URL, SHA-256, проверенный локальный путь.
- Versioned `src/domain/triggers/default-rules.ts` — заводской набор правил; main-owned settings repository хранит только пользовательскую копию/изменения и мигрирует их по stable rule id.
`Listening` является observed state только после подтверждений renderer capture и utility readiness. Main хранит desired state, но не объявляет прослушивание активным без этих acknowledgements.
### Read path
```text
navigator.mediaDevices.getUserMedia
→ AudioWorklet
→ mono PCM chunks (100300 ms)
→ transferable MessagePort with backpressure
→ utilityProcess / SpeechRecognizer adapter
→ RecognitionEvent
→ SessionController + pure domain reducer
→ versioned SessionSnapshot
→ renderer
```
### Write path
- Rules, device choice and model metadata: main-owned repository under `app.getPath('userData')` with schema validation and atomic writes.
- Index and trigger history: memory only for the current session.
- Model: temporary download, SHA-256 verification, atomic rename, then offline reuse.
- Raw audio and full transcript: never persisted in normal mode.
### Files to avoid until their phase
- `resources/models/**` and any large model binary in Git.
- `node-cpal`: microphone capture belongs to Web Audio in renderer.
- Redux/Zustand, `electron-store` and animation frameworks without a demonstrated need.
- Universal `send(channel, payload)` preload APIs.
- `ipcRenderer.invoke('acceptPcm')` as an audio transport.
- Playwright before a stable Electron flow makes its maintenance worthwhile.
- Fake implementations for the project skills named in `AGENTS.md`.
## Dependency baseline
Phase 0 pins pnpm `10.34.5` and uses Node `>=22.13 <23`, Electron `43.1.0`, Forge `7.11.2`, React `19.2.7`, TypeScript `5.9.3`, Vite `5.4.21` and Vitest `3.2.7`. Forge packages stay on one version. `.npmrc` uses `node-linker=hoisted` because Forge packaging walks physical `node_modules`; `package.json` allows build scripts only for Electron, esbuild and the pinned macOS packaging helpers.
`sherpa-onnx-node` is deliberately introduced in Phase 3, not treated as a harmless scaffold package. It is a native runtime dependency with platform packages and shared libraries. The spike must prove Vite externalization, ASAR unpack rules for both `.node` and `.dylib/.so/.dll`, runtime library lookup and a packaged-app smoke test before the dependency becomes part of the default runtime.
The first model is `sherpa-onnx-streaming-t-one-russian-2025-09-08`. Its manifest owns the expected sample rate, files, source URL, version and SHA-256. The model is not an npm dependency.
## Phase 0: reproducible application foundation
**Outcome:** a clean checkout installs, type-checks, tests, packages and opens a secure Russian-language Max Index shell on the current macOS arm64 host.
**Status:** completed on 2026-07-10; POST review passed after navigation, CSP and macOS metadata hardening.
**Files:**
- Copy `AGENTS.md` byte-for-byte from `/Users/dokril/Downloads/AGENTS.md`; preserve the already-created `PRODUCT.md` and goal documents; create `.npmrc`, `.gitignore`, `package.json`, `pnpm-lock.yaml`, `tsconfig.json`, `eslint.config.mjs`, `forge.env.d.ts`.
- Create `forge.config.ts`, `vite.main.config.ts`, `vite.preload.config.ts`, `vite.renderer.config.ts`, `vitest.config.mts`.
- Create `src/main/main.ts`, `src/preload/preload.ts`, `src/renderer/index.html`, `src/renderer/index.tsx`, `src/renderer/App.tsx`, `src/renderer/styles.css`.
- Create only the first used pure domain files for index clamping and face-level selection plus tests.
- Update `README.md` with commands, architecture status and explicit non-capabilities.
**Allowed scope:** secure `BrowserWindow`, build/release makers, static paused snapshot and original CSS placeholder. No microphone, fake listening, recognizer, persistence, tray or settings.
**Verification:**
```bash
pnpm install --frozen-lockfile
cmp /Users/dokril/Downloads/AGENTS.md AGENTS.md
pnpm lint
pnpm typecheck
pnpm test
pnpm package
pnpm start
```
**Acceptance evidence:** package succeeds on macOS arm64 and a real Electron window states that the microphone is not connected yet. Windows and Linux remain unverified.
**Captured evidence:** frozen-lockfile install, lint, strict typecheck and 14 Vitest tests pass; `pnpm package` produces `out/Max Index-darwin-arm64/Max Index.app`; the packaged process loads `file://.../app.asar/.vite/renderer/main_window/index.html` and the Paused shell was visually inspected. A forced `https://example.com` navigation remained on the allowlisted file URL. Production CSP disables connections and inline styles; `Info.plist` has ATS arbitrary loads disabled and only the microphone usage description. Source and project `AGENTS.md` share SHA-256 `c66e1d562ee2fb1a9ffbba597081508b5318b7c0770e919d73b811cac8ff9985`.
**Parallel:** renderer shell and build configuration can proceed in parallel after package/config contracts are fixed.
## Phase 1: real mock utility vertical slice
**Outcome:** `Start → mock partial/final → exactly one TriggerMatched → reaction/index → Pause → Reset` crosses the intended process boundaries.
**Files:**
- Create `src/shared/app-contract.ts` and `src/shared/recognition-contract.ts` with runtime guards.
- Create `src/domain/triggers/default-rules.ts`, `normalize-text.ts`, `match-trigger.ts`, `trigger-engine.ts`; create focused siblings under `src/domain/index/` and `src/domain/face/`, with tests next to each module.
- Create `src/main/session-controller.ts` and `src/main/utility-recognizer.ts`.
- Create `src/audio/recognition/speech-recognizer.ts`.
- Create `src/utility/recognition.ts` as a separate Forge/Vite build entry.
- Replace the static renderer snapshot with a read-only projection over typed preload commands `start`, `pause`, `reset`, `getSnapshot`, `subscribe → unsubscribe`.
**Contract details:** partial and final for one utterance share `sessionId + utteranceId`, `sequence` is monotonic, and match identity includes `ruleId + phrase position`. Cooldown is evaluated after dedupe. `SessionSnapshot.stateVersion` resolves snapshot/subscription races.
**Default-rules cutover:** Phase 1 ships the canonical defaults in code and initializes the in-memory session from them. Phase 4 introduces persisted user rules keyed by the same stable ids; there is no second editable defaults file and no renderer-owned rule source.
**Utility lifecycle:** fork only after `app.whenReady()`, use explicit dev/package artifact paths, handle ready/exit/crash/restart/shutdown, and ignore old-session events.
**Verification:** unit tests for normalization, `ё/е`, word boundaries, overlapping rules, partial/final dedupe, cooldown, clamp and Reset; integration tests for a real utility child, stale sessions and exit handling; manual mock user flow in dev and an explicitly internal, non-release packaged test artifact.
**Acceptance evidence:** one and only one visible increment for partial + final of the same utterance.
**Production cutover gate:** release packaging cannot expose or select the mock recognizer. Phase 1's packaged mock is built only under an explicit internal test flag; Phase 3 removes that flag from release configuration before sherpa becomes the default production adapter.
**Parallel:** domain reducer and renderer projection can proceed in parallel after shared contracts are fixed.
## Phase 2: microphone and bounded PCM transport
**Outcome:** Start acquires the selected microphone, Pause releases all tracks, and bounded PCM reaches the utility process without freezing renderer or main.
**Files:**
- Create `src/audio/capture/audio-capture.ts` and the AudioWorklet processor.
- Extend main/preload/utility orchestration for a transferred `MessagePort` data plane.
- Add device enumeration, capture acknowledgements and observable queue health.
**Contract details:** aggregate 128-sample worklet frames into 100300 ms mono chunks; report actual input sample rate; resample in one utility-owned location; cap the queue and emit overload state instead of growing memory; stop tracks on Pause, device switch, window shutdown and app quit.
**Verification:** Start → Listening acknowledgement, Pause track shutdown, device switch/removal, stale-session rejection, bounded queue and proof that no audio file is created.
**Acceptance evidence:** a synthetic/mock inference path consumes live microphone PCM, while UI remains responsive under a deliberately slow consumer.
**Parallel:** capture and utility queue implementation can proceed in parallel only after the MessagePort protocol is reviewed.
## Phase 3: sherpa-onnx and T-one native packaging spike
**Outcome:** the same `SpeechRecognizer` contract runs the T-one streaming model locally in both dev and packaged macOS arm64 builds.
**Files:**
- Add `sherpa-onnx-node` as a runtime dependency.
- Create `src/audio/recognition/sherpa-recognizer.ts` and model manifest/types.
- Create main-owned model download/integrity/path management.
- Update Vite externalization and Forge ASAR unpack configuration for the addon and shared libraries.
- Add a licensed short audio fixture, expected transcript assertions and packaged native-load smoke.
**Contract details:** sample rate comes from the model manifest (T-one currently expects 8 kHz); adapter synthesizes stable utterance identity from endpointing; errors distinguish recoverable runtime failure from fatal model failure; transcript text is redacted from normal logs.
**Verification:** checksum failure, interrupted download, offline reuse, real-time factor/latency sample, native addon load in packaged app and exact-phrase fixture quality.
**Acceptance evidence:** bundled app starts offline after the first verified model acquisition and recognizes the agreed Russian fixture without a cloud call.
**Parallel:** model manager and sherpa adapter may proceed in parallel after manifest and error contracts are fixed.
## Phase 4: rules, settings, tray and complete experience
**Outcome:** users can configure rules and devices, see all face states, recover from common errors and keep the app in tray.
**Files:**
- Add main-owned settings repository and schema migrations.
- Add renderer settings view, rule editor, mic test and explicitly temporary transcript preview.
- Add tray lifecycle and idempotent Start/Pause/Reset orchestration.
- Replace CSS face with an original PNG/WebP sprite sheet and a small state machine.
**Verification:** persisted rules and device fallback, all six face levels, transient reaction, reduced motion, permission/model/device errors, tray reopen, recognizer restart and no leaked listeners.
**Acceptance evidence:** a non-technical user can recover from denied permission, missing model and vanished device using one clear action per state.
**Parallel:** settings UI and original asset production can proceed in parallel after the settings and face-state contracts are fixed.
## Phase 5: release evidence
**Outcome:** signed/installable artifacts and honest support claims for Windows x64, macOS arm64/x64 and Linux x64.
**Files:** release workflow, maker settings, icons, signing/notarization configuration, artifact checksums and operator runbook.
**Verification:** clean install/upgrade/uninstall, permissions, native addon/model path, tray, microphone, offline restart and one-real-utterance flow on every claimed OS/arch.
**Acceptance evidence:** archived per-platform run with artifact hash, app version, model version, OS/arch and user-flow result. A host-only package is not cross-platform evidence.
**Parallel:** platform lanes run independently after Phase 4 is stable.
## Phase 6: system audio as a separate provider
**Outcome:** opt-in meeting audio capture for headphone scenarios without destabilizing microphone mode.
**Allowed scope:** separate Windows, macOS and Linux providers behind the audio-source boundary; explicit permissions and capability detection.
**Kill criteria:** no universal loopback promise, no hidden fallback, and microphone remains the default known-good provider.
## Project completion gate
V1 is complete only when all of the following are proven on every supported target:
1. Active listening is unmistakable and Pause releases capture resources.
2. A configured Russian phrase changes the index exactly once per utterance.
3. Index stays within `0..100`; Reset is idempotent.
4. Renderer remains responsive during slow recognition and utility restart.
5. No raw audio or full transcript is persisted by default.
6. A verified model works offline after first acquisition.
7. Installed artifacts pass the per-platform user flow; unsupported targets are not claimed.
## Primary references
- [Electron Forge Vite template](https://www.electronforge.io/templates/vite)
- [Electron Forge pnpm packaging requirement](https://www.electronforge.io/)
- [Electron utilityProcess API](https://www.electronjs.org/docs/latest/api/utility-process)
- [sherpa-onnx Node addon installation](https://k2-fsa.github.io/sherpa/onnx/javascript-api/install.html)
- [sherpa-onnx T-one model](https://k2-fsa.github.io/sherpa/onnx/pretrained_models/online-ctc/t-one-ctc-models.html)