diff options
| -rw-r--r-- | .skills/ORCHESTRATOR.md | 443 | ||||
| -rw-r--r-- | notes/retry-with-backoff-plan.md | 138 | ||||
| -rw-r--r-- | packages/kernel/src/contracts/events.ts | 1 | ||||
| -rw-r--r-- | packages/kernel/src/contracts/index.ts | 2 | ||||
| -rw-r--r-- | packages/kernel/src/contracts/runtime.ts | 44 | ||||
| -rw-r--r-- | packages/kernel/src/runtime/events.ts | 14 | ||||
| -rw-r--r-- | packages/kernel/src/runtime/index.ts | 1 | ||||
| -rw-r--r-- | packages/kernel/src/runtime/run-turn.test.ts | 535 | ||||
| -rw-r--r-- | packages/kernel/src/runtime/run-turn.ts | 167 | ||||
| -rw-r--r-- | packages/session-orchestrator/src/index.ts | 6 | ||||
| -rw-r--r-- | packages/session-orchestrator/src/orchestrator.ts | 36 | ||||
| -rw-r--r-- | packages/session-orchestrator/src/pure.test.ts | 61 | ||||
| -rw-r--r-- | packages/session-orchestrator/src/pure.ts | 47 | ||||
| -rw-r--r-- | packages/wire/src/index.ts | 26 | ||||
| -rw-r--r-- | tasks.md | 38 |
15 files changed, 1089 insertions, 470 deletions
diff --git a/.skills/ORCHESTRATOR.md b/.skills/ORCHESTRATOR.md deleted file mode 100644 index 4d5f213..0000000 --- a/.skills/ORCHESTRATOR.md +++ /dev/null @@ -1,443 +0,0 @@ -Operating manual for the dispatch orchestrator: plan topological waves of single-owner agents, summon via dispatch CLI, verify from contracts + tests (never read implementation), resolve contract gaps. Project-specific to this repo. ---- -# ORCHESTRATOR.md — how to drive this project - -> **You are the orchestrator.** You do NOT write feature code yourself. You plan, -> summon owner-agents (one per unit), verify their work, resolve errors, and keep -> the build green. This file is your complete operating manual. Read it fully -> before acting. Also read: `AGENTS.md` (the subagent constitution — you enforce -> it), `GLOSSARY.md`, `.dispatch/rules/`, `tasks.md` (live progress), and -> `notes/restructure-plan.md` (the full design + rationale; §-refs below point -> into it). - ---- - -## 0. Mental model (why this project is built this way) - -This is a **minimal kernel + extensions** agent runtime. Every feature is an -extension. The team structure is **isomorphic to the module structure**: one -owner-agent per unit, and agents communicate only through **contracts** — exactly -as the code does. Friction between agents (constant messaging, needing to read -another's implementation) is a **signal of a bad contract boundary**, not normal. - -This is a synthesis of "The AI Harness" -(https://dev.to/louaiboumediene/the-ai-harness-why-your-ai-coding-agent-is-only-as-smart-as-the-repo-you-put-it-in-cml) -with our own design. The harness layers we use: -- **Constitution** (`AGENTS.md`) — loaded by every agent. Non-obvious, project- - specific rules only. -- **Safety reflexes** (`.dispatch/rules/*.md`) — tiny, crystallized scar tissue. -- **Glossary** (`GLOSSARY.md`) — one canonical name per concept. -- **This file** — the orchestrator's workflow (the article doesn't cover this; we - added it). -- **Scoped knowledge** — rules/prompts are scoped to the *kind* of agent and the - *layer* it works in (strict for kernel/pure-core, lenient for the shell). The - article's key lesson: **scoped rules beat general rules; never write down what a - frontier model already knows** (P6). - -The 8 principles (P1–P8) live in `notes/restructure-plan.md` §1. Internalize them; -they justify every rule below. - ---- - -## 1. The golden workflow (build/modify a feature) - -1. **Plan.** Decide the unit(s); split into dependency-topological **waves** of - disjoint units, and WIDEN each wave where you can (§2a). One agent owns one unit; - it may ONLY edit its assigned files. -2. **Overlap check FIRST (anti-synonym-drift, §5.6).** Before creating anything - new, check `GLOSSARY.md` + existing code. If the request *describes* an - existing concept under a new name, steer to the canonical term (e.g. - "web-notifier" → that's a `webhook`). New term? Propose the standard/training- - baked name and **ask the user** before adding it to the glossary. Never coin a - term silently. -3. **Boundary decision is the USER's (§5.2).** "New extension vs. extend an - existing one?" — surface it to the user; never decide granularity silently. -4. **Write the prompt** to `prompts/<unit>.md` (gitignored). See §3 for the - prompt recipe. -5. **Summon the wave** via `opencode run` (see §2); disjoint units run in PARALLEL - (§2a). RE-READ `.dispatch/rules/` + the §3 scoping map before each wave — assemble - from the files, not from memory. -6. **Verify** the reports + independently re-run checks (see §4). Trust nothing - until you've re-run `typecheck`/`test`/`check` yourself. -7. **Resolve** any contract gaps / errors (see §5). -8. **Commit** the milestone with a clear message + test count. Update `tasks.md`. - ---- - -## 2. Summoning agents via `opencode run` (the harness) - -OpenCode CLI is the summon mechanism (see `notes/opencode-agents.md`). - -**Working dir:** always the repo root, -`/home/tradam/projects/dispatch/dispatch-backend` (so the agents' `lsp` tool works — -TS language server is configured globally). - -**Model:** use `opencode-go/mimo-v2.5-pro` for BUILDING agents (capable coder). -`deepseek-v4-flash` is reserved as the *app's own runtime testbench*, not for -building. - -**Canonical invocation** — assemble the prompt by CONCATENATING the standardized briefs + the -scoped rules + the per-summon TASK. The invariant guardrails live ONCE in the briefs, so -`prompts/<unit>.md` is now JUST the TASK block (§3). Do NOT use `-f` (see gotcha); ALWAYS -redirect output to a file. -```bash -cd /home/tradam/projects/dispatch/dispatch-backend && \ -opencode run --dir /home/tradam/projects/dispatch/dispatch-backend \ - -m opencode-go/mimo-v2.5-pro \ - "$(cat .dispatch/package-agent.md) -$(cat .dispatch/extension-agent.md) -$(cat .dispatch/rules/one-owner.md .dispatch/rules/isolation-over-dry.md .dispatch/rules/biome-clean.md .dispatch/rules/pure-core.md .dispatch/rules/no-internal-mocks.md .dispatch/rules/typed-handles.md) - -## TASK -$(cat prompts/<unit>.md)" \ - > reports/<unit>.run.log 2>&1 -``` -**Assembly order is fixed: package brief → extension supplement → scoped rules → TASK** -(the supplement references "the package brief above"; the briefs reference "rules inlined into -this prompt"). Rules: -- **Non-extension package?** OMIT the `.dispatch/extension-agent.md` line. -- Inline ONLY the scoped rules matching the unit's layer (the §3 map) — not every rule on every agent. -- `AGENTS.md` is auto-loaded by opencode — never `cat` it. -- The briefs already instruct the agent on ownership, visibility, verify, and the report; the - TASK block must NOT repeat any of that. - -**MANDATORY — capture output to a file, never display it.** The agent's streamed -output is enormous and will overwhelm and CRASH this harness if it lands in your -terminal. ALWAYS redirect the summon's stdout+stderr to a log file (e.g. -`> reports/<unit>.run.log 2>&1`) and do NOT echo/`cat` that log back wholesale. -You don't need the raw stream: read the agent's `reports/<unit>.md` report (and, -if you must, `grep`/`tail` the log for a specific error). Treat dumping a full run -log into context as a hard failure. - -**Run discipline (from the tool harness):** -- **Do NOT background it. Use a large timeout** (e.g. 1800000 ms = 30 min) — these - are long tasks. Backgrounding loses the stream. -- One non-backgrounded `run_shell` per summon. For PARALLEL agents on disjoint - files, launch multiple summons (the harness allows concurrent tool calls) — but - ONLY when their file sets do not overlap (single-writer rule). Log parallel runs - in `tasks.md`. - -**GOTCHAS (learned the hard way):** -- **Headless cross-`--dir` read = HANG.** An agent's Read of any file OUTSIDE its - `--dir` triggers an interactive permission prompt that CANNOT be answered headlessly - → the run wedges until aborted. This bites CROSS-REPO: a `file:` dep symlink (e.g. - `dispatch-web/node_modules/@dispatch/ui-contract` → the sibling repo) resolves OUTSIDE - `--dir`, so an agent reading the dep's source hangs. Fixes: (a) keep everything the - agent must READ inside `--dir` — ship an **in-repo reference snapshot** of a cross-repo - contract and FORBID reading `node_modules/@dispatch/*`; OR (b) set `--dir` to a parent - containing all needed paths — but then the repo's `AGENTS.md` won't auto-load (you lose - the constitution). The briefs now tell agents: never read outside your scope — if you - think you need to, REPORT it and STOP, never attempt the read. -- `-f/--file` is an ARRAY flag and greedily eats your trailing message as another - filename → "File not found". **Inline with `"$(cat prompts/X.md)"` instead.** -- A quick smoke test works: `opencode run -m opencode-go/mimo-v2.5-pro "Reply with - exactly SMOKE_OK"` should print `SMOKE_OK`. -- `opencode models` lists models; `opencode agent list` lists agent profiles; - `opencode run --help` for flags. - ---- - -## 2a. Parallel execution — WAVES - -Throughput comes from running disjoint units at once. Organise it as waves: -- **A wave = units that (a) touch DISJOINT files and (b) have no compile-time dependency - on each other** (each imports only already-built packages + existing contracts). Launch a - wave by emitting one summon per unit as CONCURRENT tool calls (§2). Later waves depend on - earlier ones; the composition root (`packages/host-bin/`) is almost always the LAST wave. -- **Pre-author the seam to widen the wave.** Because the orchestrator OWNS contracts (§6), - author the shared contract / typed handle in `packages/kernel/src/contracts/*` FIRST, then - summon the producer AND the consumer in the SAME wave against that fixed type — neither needs - the other's implementation. Authoring the contract up front is what turns a sequential - producer→consumer chain into one parallel wave (and `lsp references` on the new symbol gives - the exact consumer set to summon). -- **Also widen by removing edges:** prefer a consumer-defined handle the producer implements, - or a generic utility over a feature-specific one, so a dependency disappears entirely. -- **One writer per file, always** — even across waves. If two units would edit the same file, - they are NOT separable; merge them into one unit or sequence them. -- **After a wave:** read every report, run the §4 checks ONCE for the whole wave, commit the - milestone (update `tasks.md`), then start the next wave. Don't open a new wave before the - prior one is green. - ---- - -## 3. The per-summon `prompts/<unit>.md` is JUST the TASK block - -The invariant guardrails — single-writer directory ownership, visibility, coupling, the -engineering standard, isolated verification, and the report format — live ONCE in the -standardized briefs the summon concatenates (§2): -- **`.dispatch/package-agent.md`** — the base for EVERY package owner. -- **`.dispatch/extension-agent.md`** — the extension-only supplement (added for extension summons). - -So `prompts/<unit>.md` no longer restates any of that. It contains ONLY the **TASK**: -1. **Your package:** `packages/<name>/` — name the WHAT, not the files (the owner owns the whole - directory and decides which files to touch). -2. **The job + algorithm**, naming the specific contract types/handles involved. -3. **The specific contract file(s)** to read (e.g. `packages/kernel/src/contracts/<x>.ts`) and - any sibling public surfaces it consumes. -4. **The required test cases** (named). - -Keep it scoped (P6): state only the project-specific, non-inferable task — the briefs carry the rest. - -**`.dispatch/rules/` scoping map** — include ONLY the rows matching the unit (per §0 -"scoped rules beat general rules"); do NOT dump every rule on every agent: -- **Every agent:** `one-owner.md`, `isolation-over-dry.md`, `biome-clean.md`. -- **Kernel unit:** `kernel-purity.md` + `pure-core.md` + `no-internal-mocks.md`. -- **Pure-core unit:** `pure-core.md` + `no-internal-mocks.md`. -- **Any extension coupling via hooks/services:** `typed-handles.md`. -- **Every extension (≈ all of them — they all log):** `extension-logging.md`. Use the - injected `host.logger`/`ctx.log`; keystone: each extension self-redacts its OWN secrets - in its OWN code — NO shared redaction helper (design rationale: - `notes/observability-design.md` §9). Include this on EVERY extension summon (an - extension that never logs is a coverage gap, not an exemption). -- **Frontend units** are summoned from the SEPARATE `../dispatch-web` repo using ITS - OWN harness (`package-agent.md` + `frontend-*.md` rules) + ITS OWN scoping map — NOT - these backend rules. See that repo's `ORCHESTRATOR.md`. - -**Tell each agent it has company (parallel waves).** Add to each wave TASK: sibling units are -being built in OTHER packages right now; `tsc -b`/vitest/biome are whole-PROJECT, so if a check -reports errors OUTSIDE your package, that's concurrent WIP — ignore it and ensure YOUR files are -clean. The orchestrator's post-wave run (§4) is the source of truth. - -**Make agents IMPLEMENT, not deliberate.** A summoned owner must edit files + run its checks + -write its report in the one run. If a summon returns only a plan, re-summon (§5a). - ---- - -## 4. Verification (the orchestrator's trust protocol) - -**Plan principle (§3.6 / §5 last row):** the orchestrator confirms work from -**contracts + test results + build/diagnostics output** — that is the *designed* -trust mechanism, and it works precisely because the boundaries are testable. The -tests-at-boundaries ARE how you trust a unit without depending on its internals. - -**Stay out of implementation files (§6 Visibility).** Your trust signals are the -agent's report, the contract/surface it exposes (contracts, manifests, public -types), and the build/test/lint output you re-run yourself — NOT its implementation -code. Do NOT open an extension's implementation files — not even to "skim", -double-check, or diagnose a bug. **There is NO "conflict exception."** When X and Y -don't work together, or a unit is broken, you diagnose from the `typecheck`/`test` -output + `lsp references` on the contract + the agent's report, then **summon the -owning agent** (or a temporary multi-knowledge agent, §5) to read its own code and -fix it. You diagnose from symptoms; the agent reads the code. - -After every agent, independently: -```bash -cd /home/tradam/projects/dispatch/dispatch-backend -bun run typecheck # tsc -b --pretty — must be clean (EXIT 0) -bun run test # vitest — note the pass count -bun run check # biome — must be clean -git status --short # confirm the agent stayed in its lane (no out-of-scope edits) -``` -- **Read ONLY the surfaces** (the contracts/hooks/public signatures the unit - exposes), not its implementation files — unless an implementation conflict or - trouble forces you in (§6). The surface plus green checks is enough to trust a - unit; subtle contract mistakes show up at the boundary, which is what the - contract + boundary tests are for. -- Confirm the agent touched ONLY its assigned files (one-owner rule). -- For pure units, confirm tests use NO internal `vi.mock("@dispatch/*")`. - -**Concurrency caveat (parallel waves):** `tsc -b`/vitest/biome are whole-project, so an agent's -OWN mid-wave check can transiently see a sibling's half-written file. Don't act on a report's -out-of-package errors; YOUR post-wave run is authoritative. Re-run a suite that depends on shared -external state before trusting it — and ALWAYS sweep leaked server/collector processes between -live runs (§8 bracket trick), since a leak silently poisons the next run's counts. - ---- - -## 5. Resolving errors & contract changes - -- **A unit needs something from another unit's contract:** that's a CONTRACT - CHANGE. The owner of the contract makes it. To find every consumer, use - `lsp references` on the changed exported symbol (contracts are static TS types, - so this returns the TRUE blast radius — §5.3). Then summon the affected owners - to update. The orchestrator dispatches this fan-out; agents don't reach across. -- **Integration bug (X and Y each honor the contract but don't work together):** - no single file owns it. Summon a **temporary multi-knowledge agent** with - read/write to the 2–3 relevant files (it MAY see implementation — exception to - the visibility rule), as their temporary exclusive owner. Dispatch proactively, - or when a file-owner requests it (§5.5). -- **CR (change-request) in a report:** if it's **build/config** (root - `tsconfig.json` ref, a `package.json` dep, `.gitignore`, `bun.lock`) the - orchestrator edits it directly, then re-verifies. If it's **implementation** (a - barrel `index.ts`, a sibling conforming to a contract change), the orchestrator - **summons the owning agent** — it does NOT edit implementation itself. -- **Live API errors:** an HTTP 429 `GoUsageLimitError` is an UPSTREAM rate limit, - not a bug. The `opencode-2` key has a monthly cap; `opencode-1` is the backup. - Swap `DISPATCH_API_KEY` in `.env` (both keys are there). - ---- - -## 5a. Agent-failure recovery patterns - -- **Plan-only / "shall I proceed?" agent.** A summon sometimes returns a PLAN and STOPS without - editing (no diff, no `reports/<unit>.md`). Detect via `git status` + the missing report. - Re-summon the SAME TASK prefixed: "IMPLEMENT THIS NOW — make all edits, run the checks, write - the report; do not stop to plan or ask." Don't hand-fix its work. -- **A behaviour change reds a SIBLING's tests (test fan-out).** When a unit's new behaviour - invalidates another unit's test ASSERTIONS, those tests belong to that OTHER owner — summon it - with a focused "fix these N failing tests to match the new behaviour" TASK (state the - behaviour). The orchestrator never edits feature tests itself. (Distinct from an INTEGRATION - bug where neither side is wrong — that's the temporary multi-knowledge agent in §5.) -- **Agent strayed out of its lane.** `git status --short` after every wave; if an agent touched a - file outside its package, keep it ONLY if it's legitimately the orchestrator's lane - (contracts / build / config / harness, §6) and note it — otherwise revert + re-summon with a - tighter scope. -- **Flaky green.** A wave that passes once but leaked a server/collector or relies on shared - external state can pass for the wrong reason; sweep (§8) and re-run before committing. - ---- - -## 6. Restrictions & invariants (NEVER violate) - -- **Single-writer:** never let two agents edit the same file concurrently. -- **Kernel purity:** no I/O / no concrete feature names in `packages/kernel` - (`.dispatch/rules/kernel-purity.md`). -- **Visibility rule (§5.1):** agents see only other units' CONTRACTS, never their - implementation. A contract documents **behavior & guarantees a consumer can - rely on, not just types** (P6 applied to contracts). An agent *needing* to read - another unit's code is a signal that contract is underspecified — fix the - contract, don't grant code access. (Exception: the temporary multi-knowledge - integration agent, §5 / ORCHESTRATOR §5, which MAY read implementation.) -- **The orchestrator NEVER reads or edits implementation.** You read ONLY contracts - (`packages/kernel/src/contracts/*`) + surfaces (manifests, public signatures) + - diagnostics (`typecheck`/`test` output, `lsp references` on contract symbols) + - agent reports. Do NOT open implementation `.ts` files (feature logic, tests, - composition roots) — not even during a bug. Clean context = level-headed - decisions; the subagents do the implementation. -- **What the orchestrator MAY edit directly:** (a) **contracts** - (`packages/kernel/src/contracts/*`); (b) **build wiring + config** (root/package - `tsconfig.json`, `package.json` deps, project refs, `.gitignore`, `bun.lock`); - (c) **harness/docs** (`ORCHESTRATOR.md`, `AGENTS.md`, `GLOSSARY.md`, - `.dispatch/rules/`, `notes/`, `tasks.md`, `prompts/`, `reports/`). Everything else - — all executable implementation `.ts`, including tests and composition roots like - `host-bin/src/main.ts` — changes ONLY by summoning the owning agent. -- **Roadblock → surface to the user.** If a needed change doesn't fit the above - (ambiguous ownership, a design question, a stuck agent), stop and ask rather than - reaching into implementation. -- **Subagents inherit this restriction.** Every prompt you write must instruct the - agent to read ONLY the surfaces (contracts/hooks) of OTHER units, with the sole - exception that it MAY read the implementation files of the task/extension it is - assigned to. It must not go spelunking through sibling units' implementations. -- **`onAny` is the ONLY allowed dynamic hook subscription** (observability/logging - firehose). All other cross-extension coupling is typed-symbol anchored (§5.4). -- **Contracts are static TYPES; loading is dynamic** (manifests via host). This - split is load-bearing — it's what makes `lsp references` fan-out work. -- **Full fidelity:** every core feature is a real extension with a manifest, - loaded through the host. Do NOT hand-wire imports to shortcut the extension - model — that defeats the point. -- **Asymmetric testing:** strict (zero internal mocks, high coverage) on - kernel/pure-core; lenient (thin integration tests) on the shell. -- **Destructive git ops:** be extremely careful. Back up irreplaceable files - (e.g. `notes/restructure-plan.md`) to `/tmp` before any `git reset --hard` / - `git clean`. `.env` is gitignored — preserve it. -- **Write things up before pivoting topics.** Keep `tasks.md` current in real - time. Don't leave decisions only in chat context (it can be lost). - ---- - -## 7. Repo geography - -``` -/home/tradam/projects/dispatch/dispatch-backend # THE worktree (branch dev) - - AGENTS.md the subagent constitution (auto-loaded by opencode; you enforce it) - ORCHESTRATOR.md the orchestrator's operating manual (this file) - GLOSSARY.md canonical vocabulary + aliases-to-avoid (human-gated) - tasks.md live progress checklist / milestone log - README.md deployment, CLI usage, extension/package tables - - .dispatch/ - package-agent.md base owner-agent brief (every summon) - extension-agent.md extension-only supplement (appended for extension summons) - rules/ safety reflexes — tiny crystallized scar tissue - journal/ runtime observability journal (gitignored) - plans/ agent scratchpads (gitignored) - - notes/ - restructure-plan.md the full architecture design + rationale (P1–P8; §-refs) - observability-design.md logging/spans/collector/trace-store design (Phase A–B) - cli-design.md CLI design decisions + unit plan (built; §3 = settled decisions) - frontend-design.md future web frontend design (IDEATION; separate repo) - opencode-agents.md notes on summoning agents via the opencode CLI - - prompts/ (gitignored — orchestrator→agent TASK blocks) - reports/ (gitignored — agent→orchestrator reports) - .env (gitignored — DISPATCH_API_KEY, DISPATCH_BASE_URL, DISPATCH_MODEL, BACKEND_PORT) - - packages/ - kernel/ contracts (ABI), bus, runtime (runTurn), host - wire/ types-only wire ABI (AgentEvent + conversation model + Usage); kernel + - transport-contract re-export it so clients consume the wire w/o the kernel runtime - transport-contract/ types-only HTTP API contract (CLI + future web + server share it) - ui-contract/ types-only surface ABI (frontend-agnostic; web + CLI render it) - storage-sqlite/ conversation-store/ auth-apikey/ provider-openai-compat/ - credential-store/ named credentials + model catalog (resolve / listCatalog) - session-orchestrator/ transport-http/ (core extensions) - tool-read-file/ standard tool extension (read_file; cwd-aware) - journal-sink/ trace-store/ observability-collector/ trace-replay/ (observability) - cli/ bundled one-shot terminal client (HTTP client of transport-contract) - host-bin/ composition root (boot + Bun.serve + collector supervisor) -``` - -The genesis commit deleted all prior source; we rebuilt from scratch. The OLD -project lives at `/home/tradam/projects/dispatch/dispatch-source` (reference only -— do not edit). - -The **web frontend is a SEPARATE repo** at `/home/tradam/projects/dispatch/dispatch-web` -(own git, own harness — its own `AGENTS.md`/`ORCHESTRATOR.md`/`GLOSSARY.md`/`.dispatch/`). -It consumes `packages/ui-contract` + the wire types as a pinned `file:` dependency. -`lsp references` does NOT span the two repos, so cross-repo contract changes are -**couriered via the user** (see the FE `ORCHESTRATOR.md` §5). Design + plan: -`notes/frontend-design.md`. Do NOT edit the FE repo from here. - ---- - -## 8. Current status & how to run - -See `tasks.md` for the live checklist. As of MVP completion: -- Kernel + 6 core extensions + host-bin DONE. 178 tests pass; typecheck + biome - clean. -- **MVP verified live:** multi-turn curl against OpenCode Go flash works - (`conversationId` threads history). - -**Boot + smoke test:** -```bash -cd /home/tradam/projects/dispatch/dispatch-backend -KEY1=$(grep DISPATCH_API_KEY_OPENCODE1 .env | cut -d= -f2) -PORT=4567 DISPATCH_API_KEY="$KEY1" bun packages/host-bin/src/main.ts # boots server -# in another shell: -curl -s -X POST localhost:4567/chat -H 'content-type: application/json' \ - -d '{"conversationId":"c1","message":"Say hello in 3 words."}' -``` -Note the chat field is **`conversationId`** (threads multi-turn), not `tabId`. - -**Live validation & process cleanup — the `[x]` bracket trick (scar tissue).** When -you live-validate you background the app (`bun packages/host-bin/src/main.ts &`), and it -now spawns a child **observability collector** process. To list or kill those, ALWAYS -use the bracket trick in `ps`/`pgrep`/`pkill` patterns: -```bash -ps -eo pid,args | grep '[o]bservability-collector/src/main' # list (won't self-match) -pkill -9 -f '[h]ost-bin/src/main.ts' # kill the app -pkill -9 -f '[o]bservability-collector/src/main' # kill the collector -``` -**Why it matters:** a plain `pkill -f 'host-bin/src/main.ts'` matches its OWN command -line and kills the parent shell → the tool call prints NOTHING and times out, looking -exactly like a wedged session. `[h]ost-bin` matches the target "host-bin" while the -literal pattern `[h]ost-bin` does not match itself. ALWAYS clean up the backgrounded app -+ its spawned collector after each live run — leaked processes pollute the next run's -counts (this is precisely what made a correct supervisor look like it spawned 3 -collectors and left 2 behind). - -**Live boot-probe in ONE command WILL hit the tool timeout — that is NOT failure (scar tissue).** -A single bash command that boots the app (even detached via `setsid … & disown`), sleeps, runs a -probe, then kills it will still run to the tool's timeout: the tool waits on the spawned -server/collector session. The probe already ran — **read the probe's printed `RESULT: OK/FAIL` -line as the signal**, ignore the timeout, then run a SEPARATE `pkill` (bracket-trick) + `ps` -cleanup command (it returns immediately and confirms no leaks). Don't try to make the boot+probe -command "return cleanly" — it won't. (For a frontend-agnostic surface, the probe is a tiny -`bun` WebSocket client that asserts `catalog → subscribe → surface`.) - -**Next suggested work** (post-MVP, see `tasks.md` "Open items"): wire -auth→provider properly (auth-apikey is currently vestigial), then add the first -TOOL extension to exercise the dispatch loop (turns currently run with `tools: -[]`). diff --git a/notes/retry-with-backoff-plan.md b/notes/retry-with-backoff-plan.md new file mode 100644 index 0000000..99f6f5d --- /dev/null +++ b/notes/retry-with-backoff-plan.md @@ -0,0 +1,138 @@ +# Plan — Retry-with-backoff on retryable provider errors (FINALIZED) + +**Goal:** When the upstream LLM API returns a retryable error (e.g. "server +overloaded"), retry the request with a stepped backoff, visibly, until the +budget is exhausted. + +## The error (from the prod DB) — detection is already done + +- **HTTP 429** (46×) and **HTTP 502** (1×), **no 503s**. +- Body: `{"error":{"type":"overloaded_error","message":"The service is temporarily overloaded. Please retry."}}` +- `packages/openai-stream/src/stream.ts:201` **already sets** + `retryable: response.status >= 500 || response.status === 429` on the error + event, and `ProviderErrorEvent` (`kernel/contracts/provider.ts:72`) **already + declares `retryable?: boolean`**. The kernel's `processEvent` just ignores it. +- The error is **emitted (not thrown) and before any content** → retrying + `provider.stream()` is safe (no partial chunks to roll back). + +## Decision 1 — the backoff schedule + +`5s, 10s, 30s, 60s, 5m, 10m, 15m, 30m`, then **repeat 30m** until **8h of +cumulative retry-wait** is reached, then give up (emit the final error + seal). + +Pure function of the attempt index (0 = first retry): +```ts +const SCHEDULE_MS = [5_000, 10_000, 30_000, 60_000, 300_000, 600_000, 900_000, 1_800_000]; +const TAIL_MS = 1_800_000; // 30m +const BUDGET_MS = 8 * 60 * 60 * 1000; // 8h + +// pure, deterministic, no I/O +function delayFor(attempt: number): number | undefined { + const delay = attempt < SCHEDULE_MS.length ? SCHEDULE_MS[attempt] : TAIL_MS; + if (cumulativeSleepMs(attempt) > BUDGET_MS) return undefined; // over budget → stop + return delay; +} +``` +- `cumulativeSleepMs(attempt)` = sum of delay[0..attempt]; head (8 steps) sums to + 3,705s, then +1,800s per extra step. 8h (28,800s) is reached at attempt ~21 + → ~21 retries, ~7h32m of sleeping, then give up. +- Budget = cumulative *scheduled sleep* (pure/testable). If you prefer wall-clock + since first error, it switches to using the injected `now` — easy change. + +## Decision 2 — visible (yellow system-message warning) + 5d3f handoff + +Add a new **transient** `AgentEvent` variant (emitted to the frontend, NOT +persisted into the model's message history — so it never pollutes the prompt): + +```ts +// @dispatch/wire (AgentEvent union gains this member) +export interface TurnProviderRetryEvent { + readonly type: "provider-retry"; + readonly conversationId: string; + readonly turnId: string; + /** 0-based: this is the Nth retry about to happen. */ + readonly attempt: number; + /** ms the client should expect to wait before the retry fires. */ + readonly delayMs: number; + /** The endpoint's error verbatim (e.g. "HTTP 429: {…overloaded_error…}"). */ + readonly message: string; + /** The HTTP code when known (e.g. "429"). */ + readonly code?: string; +} +``` +- Emitted once per scheduled retry, BEFORE the sleep, so the UI shows + "⚠ Server overloaded — retrying in 5s…" immediately. +- When retries are exhausted (8h), the existing `error` event is emitted (as + today) and the turn seals — so the final failure is still a persisted error. + +**Frontend handoff to 5d3f:** render `provider-retry` as a yellow warning +system-message bubble showing `message` (+ `code`), with the countdown. (I do +the backend; 5d3f does the renderer — handoff via dispatch CLI.) + +## Decision 3 — retry ANY retryable error + +Retry trigger (both paths), **only when no content has been emitted yet** +(the safety invariant — never duplicate partial output): + +- **Emitted** `error` ProviderEvent with `retryable === true` → retry. (429/502/5xx + network fetch errors — all pre-content.) +- **Thrown** error (mid-stream, caught in `executeStep`'s `catch`) → treated as **retryable-by-default when pre-content** (most mid-stream throws are transient network/SSE issues). A thrown error after content is emitted is NOT retried (can't safely). + +So "if it's retryable, retry it" = the `retryable` flag drives emitted errors; +thrown errors default to retryable when nothing was streamed yet. Non-retryable +emitted errors (`retryable: false`/absent) end the step as today. + +## Architecture — kernel provides the HOOK, shell provides POLICY + I/O + +(Constitution: kernel touches no I/O; effects injected; decision pure.) + +### Kernel contract (`kernel/src/contracts/runtime.ts`) — add to `RunTurnInput`: +```ts +export interface RetryStrategy { + /** Pure: attempt → delay ms, or undefined to stop (budget exhausted). */ + readonly delayFor: (attempt: number) => number | undefined; + /** Injected effect: actually sleep. Kernel imports no timer. Abortable. */ + readonly sleep: (ms: number, signal: AbortSignal) => Promise<void>; +} +export interface RunTurnInput { + // …existing… + /** Optional injected retry. Omit = no retry (backward-compatible). */ + readonly retry?: RetryStrategy; +} +``` + +### Kernel loop (`kernel/src/runtime/run-turn.ts`, `executeStep`): +Wrap stream consumption in a retry loop: +- track `hadContent` (any text/reasoning/tool-call/usage seen); +- on a retryable error (emitted `retryable:true` OR thrown) with `!hadContent`: + - `delay = retry.delayFor(attempt)`; if `undefined` → give up (emit the + suppressed error, end step); + - else emit `providerRetryEvent(attempt, delay, message, code)`, `await + retry.sleep(delay, signal)`, `attempt++`, re-call `provider.stream()`; +- on abort during sleep → reject, seal turn `aborted` (existing flow). + +### Shell wiring (`session-orchestrator/src/orchestrator.ts`): +- Provide the concrete `RetryStrategy`: `delayFor` = the schedule + 8h budget + above; `sleep` = abortable `setTimeout`-based promise. +- Pass `retry` into the `RunTurnInput` it builds (line 589). + +## Build breakdown by unit (execution) + +| Unit (owner) | Change | +|---|---| +| `@dispatch/wire` | add `TurnProviderRetryEvent` to `AgentEvent` union | +| `kernel` contracts | add `RetryStrategy` + `retry?` on `RunTurnInput` | +| `kernel` events.ts | `providerRetryEvent(...)` constructor | +| `kernel` run-turn.ts | retry loop in `executeStep` (the core logic) | +| `kernel` run-turn.test.ts | pure tests: fake `sleep` + pure `delayFor`; assert schedule, no-after-content retry, give-up emits error, abort-during-sleep | +| `session-orchestrator` | wire concrete schedule + real `setTimeout` sleep | +| `transport-ws` | if it has an exhaustive `switch(event.type)`, add the `provider-retry` case | +| `transport-http` (mine) | **no change** — `serializeEventLine` is generic `JSON.stringify` | +| frontend (5d3f) | render `provider-retry` as a yellow warning system message | + +## Open items +- **8h budget = cumulative scheduled sleep** (pure). Confirm OK vs wall-clock. +- **Thrown errors default retryable-when-pre-content.** Confirm (vs only the + flagged emitted path). +- **Execution mode:** this spans kernel + wire + orchestrator (outside my + transport-http unit). Build it directly across units, or dispatch each slice + to its unit owner via the dispatch CLI? diff --git a/packages/kernel/src/contracts/events.ts b/packages/kernel/src/contracts/events.ts index 6c9652d..dca34c2 100644 --- a/packages/kernel/src/contracts/events.ts +++ b/packages/kernel/src/contracts/events.ts @@ -11,6 +11,7 @@ export type { TurnDoneEvent, TurnErrorEvent, TurnInputEvent, + TurnProviderRetryEvent, TurnReasoningDeltaEvent, TurnSealedEvent, TurnStartEvent, diff --git a/packages/kernel/src/contracts/index.ts b/packages/kernel/src/contracts/index.ts index c67607b..f3e5bca 100644 --- a/packages/kernel/src/contracts/index.ts +++ b/packages/kernel/src/contracts/index.ts @@ -40,6 +40,7 @@ export type { TurnDoneEvent, TurnErrorEvent, TurnInputEvent, + TurnProviderRetryEvent, TurnReasoningDeltaEvent, TurnSealedEvent, TurnStartEvent, @@ -109,6 +110,7 @@ export type { export type { EventEmitter, FinishReason, + RetryStrategy, RunTurnInput, RunTurnResult, } from "./runtime.js"; diff --git a/packages/kernel/src/contracts/runtime.ts b/packages/kernel/src/contracts/runtime.ts index 02fc446..8376e42 100644 --- a/packages/kernel/src/contracts/runtime.ts +++ b/packages/kernel/src/contracts/runtime.ts @@ -129,6 +129,22 @@ export interface RunTurnInput { * double-persist them. */ readonly onStepComplete?: (messages: readonly ChatMessage[]) => Promise<void> | void; + + /** + * Optional injected retry strategy for retryable provider errors (e.g. HTTP + * 429 / 5xx "overloaded"). When omitted, a retryable error ends the step + * exactly as before (backward-compatible). When provided, the runtime wraps + * `provider.stream()` consumption in a retry loop: on a retryable error + * (an emitted `error` ProviderEvent with `retryable === true`, OR a thrown + * error) — ONLY when no content was emitted yet this step (the safety + * invariant — never duplicate partial output) — it asks `retry.delayFor` + * for a delay, emits a transient `provider-retry` AgentEvent, sleeps via the + * injected `retry.sleep` (abortable), and re-calls `provider.stream()`. + * + * Injected (not ambient): the kernel imports no timer and owns no schedule. + * Mirrors the `now`/`logger` injection pattern — optional + backward-compatible. + */ + readonly retry?: RetryStrategy; } /** @@ -145,3 +161,31 @@ export interface RunTurnResult { /** Why the turn ended. */ readonly finishReason: FinishReason; } + +/** + * Injected retry strategy for retryable provider errors (e.g. HTTP 429 / 5xx). + * + * The kernel provides the HOOK (this contract + the retry loop in `runTurn`); + * the shell (session-orchestrator) provides the POLICY (the concrete schedule) + * and the I/O (the actual sleep). The kernel imports no timer — `sleep` is an + * injected effect so the runtime stays pure and deterministic in tests. + * + * Retries are ONLY attempted when NO content was emitted yet this step (the + * safety invariant — never duplicate partial output). When omitted on + * `RunTurnInput`, no retry happens (backward-compatible: a retryable error ends + * the step exactly as before). + */ +export interface RetryStrategy { + /** + * Pure, deterministic decision: given the 0-based attempt index, return the + * delay in ms to sleep before the next retry, or `undefined` to stop (budget + * exhausted). No I/O, no clock — fully testable. + */ + readonly delayFor: (attempt: number) => number | undefined; + /** + * Injected effect: actually sleep for the given ms. Must honor the abort + * signal — reject when aborted so the turn seals `aborted`. The kernel + * imports no timer; the shell provides a `setTimeout`-based implementation. + */ + readonly sleep: (ms: number, signal: AbortSignal) => Promise<void>; +} diff --git a/packages/kernel/src/runtime/events.ts b/packages/kernel/src/runtime/events.ts index b194577..5805e28 100644 --- a/packages/kernel/src/runtime/events.ts +++ b/packages/kernel/src/runtime/events.ts @@ -164,3 +164,17 @@ export function errorEvent( } return { type: "error", conversationId, turnId, message }; } + +export function providerRetryEvent( + conversationId: string, + turnId: string, + attempt: number, + delayMs: number, + message: string, + code?: string, +): AgentEvent { + if (code !== undefined) { + return { type: "provider-retry", conversationId, turnId, attempt, delayMs, message, code }; + } + return { type: "provider-retry", conversationId, turnId, attempt, delayMs, message }; +} diff --git a/packages/kernel/src/runtime/index.ts b/packages/kernel/src/runtime/index.ts index e1156e3..e0dd656 100644 --- a/packages/kernel/src/runtime/index.ts +++ b/packages/kernel/src/runtime/index.ts @@ -2,6 +2,7 @@ export type { StepDispatcher } from "./dispatch.js"; export { createStepDispatcher, executeToolCall } from "./dispatch.js"; export { errorEvent, + providerRetryEvent, reasoningDeltaEvent, textDeltaEvent, toolCallEvent, diff --git a/packages/kernel/src/runtime/run-turn.test.ts b/packages/kernel/src/runtime/run-turn.test.ts index 0d4c59d..dba9d80 100644 --- a/packages/kernel/src/runtime/run-turn.test.ts +++ b/packages/kernel/src/runtime/run-turn.test.ts @@ -2821,4 +2821,539 @@ describe("runTurn", () => { expect(drainCallCount).toBe(MAX_STEPS - 1); }); }); + + // ── Retry with backoff ────────────────────────────────────────────────── + // + // PURE tests: a fake `sleep` (records calls, resolves instantly, can abort + // on a chosen call) + a pure `delayFor` (the canonical schedule + 8h budget). + // A stub `ProviderContract` whose `stream` yields a retryable error N times + // then a finish. ZERO mocks of `@dispatch/*` modules — effects injected. + + /** The canonical backoff schedule (matches the orchestrator's concrete strategy). */ + const RETRY_SCHEDULE_MS = [5_000, 10_000, 30_000, 60_000, 300_000, 600_000, 900_000, 1_800_000]; + const RETRY_TAIL_MS = 1_800_000; // 30m + const RETRY_BUDGET_MS = 8 * 60 * 60 * 1000; // 8h + + /** Cumulative scheduled sleep through `attempt` (sum of delay[0..attempt]). */ + function cumulativeSleepMs(attempt: number): number { + let sum = 0; + for (let i = 0; i <= attempt; i++) { + sum += i < RETRY_SCHEDULE_MS.length ? RETRY_SCHEDULE_MS[i] : RETRY_TAIL_MS; + } + return sum; + } + + /** Pure, deterministic delay decision (no I/O, no clock). */ + function delayFor(attempt: number): number | undefined { + const delay = attempt < RETRY_SCHEDULE_MS.length ? RETRY_SCHEDULE_MS[attempt] : RETRY_TAIL_MS; + if (cumulativeSleepMs(attempt) > RETRY_BUDGET_MS) return undefined; // over budget → stop + return delay; + } + + /** The full schedule delayFor would emit (until budget exhausted). */ + function fullSchedule(): number[] { + const result: number[] = []; + let attempt = 0; + while (true) { + const delay = delayFor(attempt); + if (delay === undefined) break; + result.push(delay); + attempt++; + } + return result; + } + + /** + * Fake, controllable `sleep`: records every call's delay, resolves + * instantly (no real waiting), and can abort the controller on a chosen + * 1-based call index to simulate "abort during sleep". + */ + function createFakeSleep(controller: AbortController): { + sleep: (ms: number, signal: AbortSignal) => Promise<void>; + calls: number[]; + abortOnCall: (n: number) => void; + } { + const calls: number[] = []; + let abortAt: number | undefined; + const sleep = async (ms: number, _signal: AbortSignal): Promise<void> => { + calls.push(ms); + if (abortAt !== undefined && calls.length === abortAt) { + controller.abort(); + throw new Error("aborted"); + } + // Otherwise resolve instantly (no real waiting). + }; + return { + sleep, + calls, + abortOnCall: (n: number) => { + abortAt = n; + }, + }; + } + + /** A provider that yields a retryable error `errorCount` times, then success. */ + function createRetryingProvider(opts: { + errorCount: number; + error?: { message: string; code?: string; retryable?: boolean }; + success?: ProviderEvent[]; + }): { provider: ProviderContract; streamCalls: { value: number } } { + const streamCalls = { value: 0 }; + const error: ProviderEvent = { + type: "error", + message: opts.error?.message ?? "overloaded", + ...(opts.error?.code !== undefined ? { code: opts.error.code } : {}), + ...(opts.error?.retryable !== undefined ? { retryable: opts.error.retryable } : {}), + }; + const success = opts.success ?? [ + { type: "text-delta", delta: "hi" }, + { type: "finish", reason: "stop" }, + ]; + const provider: ProviderContract = { + id: "fake", + stream() { + const idx = streamCalls.value++; + return (async function* () { + if (idx < opts.errorCount) { + yield error; + return; + } + for (const event of success) yield event; + })(); + }, + }; + return { provider, streamCalls }; + } + + describe("retry with backoff", () => { + it("retries a retryable emitted error on schedule then succeeds", async () => { + const { provider } = createRetryingProvider({ + errorCount: 3, + error: { message: "HTTP 429: overloaded", code: "429", retryable: true }, + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(result.finishReason).toBe("stop"); + // 3 retries: 5s, 10s, 30s. + expect(fake.calls).toEqual([5_000, 10_000, 30_000]); + // 3 provider-retry events (one per sleep), then the successful text. + const retryEvents = events.filter((e) => e.type === "provider-retry"); + expect(retryEvents).toHaveLength(3); + if (retryEvents[0]?.type === "provider-retry") { + expect(retryEvents[0].attempt).toBe(0); + expect(retryEvents[0].delayMs).toBe(5_000); + expect(retryEvents[0].message).toBe("HTTP 429: overloaded"); + expect(retryEvents[0].code).toBe("429"); + expect(retryEvents[0].conversationId).toBe("conv-1"); + expect(retryEvents[0].turnId).toBe("turn-1"); + } + if (retryEvents[1]?.type === "provider-retry") { + expect(retryEvents[1].attempt).toBe(1); + expect(retryEvents[1].delayMs).toBe(10_000); + } + if (retryEvents[2]?.type === "provider-retry") { + expect(retryEvents[2].attempt).toBe(2); + expect(retryEvents[2].delayMs).toBe(30_000); + } + // The error was suppressed (no error event emitted — retry succeeded). + expect(events.filter((e) => e.type === "error")).toHaveLength(0); + // The successful content still streams. + const deltas = events.filter((e) => e.type === "text-delta"); + expect(deltas).toHaveLength(1); + }); + + it("sleep is called with the full schedule [5s,10s,30s,60s,5m,10m,15m,30m,30m…]", async () => { + // Provider errors forever → retries until budget exhausted → gives up. + const { provider } = createRetryingProvider({ + errorCount: Number.POSITIVE_INFINITY, + error: { message: "overloaded", code: "429", retryable: true }, + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + // Budget exhausted → give up → error. + expect(result.finishReason).toBe("error"); + + // The sleep schedule matches the pure delayFor output exactly. + expect(fake.calls).toEqual(fullSchedule()); + + // Head of the schedule (the 8 stepped delays). + expect(fake.calls.slice(0, 8)).toEqual([ + 5_000, 10_000, 30_000, 60_000, 300_000, 600_000, 900_000, 1_800_000, + ]); + // Tail repeats 30m. + expect(fake.calls[8]).toBe(1_800_000); + expect(fake.calls.at(-1)).toBe(1_800_000); + + // 8h cumulative budget cap: head (3705s) + 13×30m = ~7h31m, then stop. + // 21 retries (attempts 0..20), then delayFor(21) → undefined → give up. + expect(fake.calls).toHaveLength(21); + const totalSlept = fake.calls.reduce((a, b) => a + b, 0); + expect(totalSlept).toBeLessThanOrEqual(RETRY_BUDGET_MS); + expect(totalSlept).toBe(3_705_000 + 13 * 1_800_000); // 27_105_000 + + // One provider-retry per sleep, plus a final error (give-up). + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(21); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + const errEvt = events.find((e) => e.type === "error"); + if (errEvt?.type === "error") { + expect(errEvt.message).toBe("overloaded"); + expect(errEvt.code).toBe("429"); + } + }); + + it("does NOT retry after content was emitted (safety invariant)", async () => { + // Provider yields text (content) THEN a retryable error. Because content + // was emitted, retrying is unsafe (would duplicate partial output). + let callCount = 0; + const provider: ProviderContract = { + id: "fake", + stream() { + callCount++; + return (async function* () { + yield { type: "text-delta", delta: "partial" } as ProviderEvent; + yield { + type: "error", + message: "overloaded", + code: "429", + retryable: true, + } as ProviderEvent; + })(); + }, + }; + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + // No retries: stream called exactly once. + expect(callCount).toBe(1); + expect(fake.calls).toHaveLength(0); + // The error is emitted (give-up) and partial content preserved. + expect(result.finishReason).toBe("error"); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(0); + expect(events.filter((e) => e.type === "text-delta")).toHaveLength(1); + }); + + it("does NOT retry a non-retryable emitted error (retryable: false)", async () => { + const { provider, streamCalls } = createRetryingProvider({ + errorCount: 1, + error: { message: "bad request", code: "400", retryable: false }, + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(streamCalls.value).toBe(1); // no retry + expect(fake.calls).toHaveLength(0); + expect(result.finishReason).toBe("error"); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(0); + }); + + it("does NOT retry a non-retryable emitted error (retryable absent)", async () => { + const { provider, streamCalls } = createRetryingProvider({ + errorCount: 1, + error: { message: "bad request", code: "400" }, // no retryable field + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(streamCalls.value).toBe(1); // no retry + expect(fake.calls).toHaveLength(0); + expect(result.finishReason).toBe("error"); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + }); + + it("give-up emits the final error when budget is exhausted", async () => { + // Custom delayFor that allows exactly 1 retry then stops. + const shortDelayFor = (attempt: number): number | undefined => + attempt === 0 ? 100 : undefined; + const { provider } = createRetryingProvider({ + errorCount: Number.POSITIVE_INFINITY, + error: { message: "overloaded", code: "429", retryable: true }, + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor: shortDelayFor, sleep: fake.sleep }, + }); + + expect(result.finishReason).toBe("error"); + expect(fake.calls).toEqual([100]); // one retry, then give up + // One provider-retry (attempt 0), then the final error. + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(1); + const errs = events.filter((e) => e.type === "error"); + expect(errs).toHaveLength(1); + if (errs[0]?.type === "error") { + expect(errs[0].message).toBe("overloaded"); + expect(errs[0].code).toBe("429"); + } + }); + + it("abort during sleep seals the turn aborted", async () => { + const { provider } = createRetryingProvider({ + errorCount: Number.POSITIVE_INFINITY, + error: { message: "overloaded", code: "429", retryable: true }, + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + fake.abortOnCall(2); // abort on the 2nd sleep + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(result.finishReason).toBe("aborted"); + // Two sleeps attempted; the 2nd aborted. + expect(fake.calls).toHaveLength(2); + // No terminal error emitted (it was an abort, not a give-up). + expect(events.filter((e) => e.type === "error")).toHaveLength(0); + // One provider-retry before the aborted sleep (attempt 0). + const retries = events.filter((e) => e.type === "provider-retry"); + expect(retries).toHaveLength(2); + // The done event carries reason "aborted". + const done = events.find((e) => e.type === "done"); + if (done?.type === "done") { + expect(done.reason).toBe("aborted"); + } + }); + + it("omitting retry keeps the pre-retry behavior (backward-compatible)", async () => { + // A retryable error with no retry configured → ends the step as today. + const { provider, streamCalls } = createRetryingProvider({ + errorCount: 1, + error: { message: "overloaded", code: "429", retryable: true }, + }); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + // no retry field + }); + + expect(streamCalls.value).toBe(1); // no retry + expect(result.finishReason).toBe("error"); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(0); + }); + + it("retries a THROWN error (retryable-by-default when pre-content)", async () => { + // A thrown error (no retryable flag) before content is retried. + let callCount = 0; + const provider: ProviderContract = { + id: "fake", + stream() { + callCount++; + return (async function* () { + if (callCount <= 2) { + throw new Error("network blip"); + } + yield { type: "text-delta", delta: "hi" } as ProviderEvent; + yield { type: "finish", reason: "stop" } as ProviderEvent; + })(); + }, + }; + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(callCount).toBe(3); // 2 throws retried, 3rd succeeds + expect(fake.calls).toEqual([5_000, 10_000]); + expect(result.finishReason).toBe("stop"); + expect(events.filter((e) => e.type === "provider-retry")).toHaveLength(2); + // Thrown errors have no code. + if (events[0]?.type === "provider-retry") { + expect(events[0].code).toBeUndefined(); + expect(events[0].message).toBe("network blip"); + } + expect(events.filter((e) => e.type === "error")).toHaveLength(0); + }); + + it("does NOT retry a thrown error after content was emitted", async () => { + let callCount = 0; + const provider: ProviderContract = { + id: "fake", + stream() { + callCount++; + return (async function* () { + yield { type: "text-delta", delta: "partial" } as ProviderEvent; + throw new Error("network blip"); + })(); + }, + }; + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + const result = await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + expect(callCount).toBe(1); + expect(fake.calls).toHaveLength(0); + expect(result.finishReason).toBe("error"); + expect(events.filter((e) => e.type === "error")).toHaveLength(1); + expect(events.filter((e) => e.type === "text-delta")).toHaveLength(1); + }); + + it("provider-retry events interleave correctly: error → retry-event → sleep → retry", async () => { + // Verify ordering: each provider-retry event comes BEFORE its sleep, + // and the successful content comes only after the last retry. + const { provider } = createRetryingProvider({ + errorCount: 2, + error: { message: "overloaded", code: "429", retryable: true }, + success: [ + { type: "text-delta", delta: "ok" }, + { type: "finish", reason: "stop" }, + ], + }); + const controller = new AbortController(); + const fake = createFakeSleep(controller); + + const { events, emit } = createCollectingEmit(); + + await runTurn({ + provider, + messages: [userMessage], + tools: [], + dispatch: { maxConcurrent: 1, eager: false }, + conversationId: "conv-1", + turnId: "turn-1", + emit, + signal: controller.signal, + retry: { delayFor, sleep: fake.sleep }, + }); + + const types = events.map((e) => e.type); + // turn-start, provider-retry(0), provider-retry(1), text-delta, step-complete, done + expect(types[0]).toBe("turn-start"); + const firstRetryIdx = types.indexOf("provider-retry"); + const textIdx = types.indexOf("text-delta"); + expect(firstRetryIdx).toBeGreaterThan(0); + expect(textIdx).toBeGreaterThan(firstRetryIdx); + // Both retries precede the text. + const retryCount = types.filter((t) => t === "provider-retry").length; + expect(retryCount).toBe(2); + }); + }); }); diff --git a/packages/kernel/src/runtime/run-turn.ts b/packages/kernel/src/runtime/run-turn.ts index f5d80d3..08f8459 100644 --- a/packages/kernel/src/runtime/run-turn.ts +++ b/packages/kernel/src/runtime/run-turn.ts @@ -6,12 +6,18 @@ import type { ProviderStreamOptions, Usage, } from "../contracts/provider.js"; -import type { EventEmitter, RunTurnInput, RunTurnResult } from "../contracts/runtime.js"; +import type { + EventEmitter, + RetryStrategy, + RunTurnInput, + RunTurnResult, +} from "../contracts/runtime.js"; import type { ToolCall, ToolContract } from "../contracts/tool.js"; import { createStepDispatcher, type StepDispatcher } from "./dispatch.js"; import { doneEvent, errorEvent, + providerRetryEvent, reasoningDeltaEvent, stepCompleteEvent, textDeltaEvent, @@ -120,6 +126,8 @@ interface StepContext { readonly now: (() => number) | undefined; /** Per-turn provider options (model, systemPrompt, …) threaded to stream(). */ readonly providerOpts: ProviderStreamOptions | undefined; + /** Optional injected retry strategy (omit = no retry, backward-compatible). */ + readonly retry: RetryStrategy | undefined; } interface TimingState { @@ -249,12 +257,10 @@ function processEvent( case "finish": break; case "error": - if (event.code !== undefined) { - chunks.push({ type: "error", message: event.message, code: event.code }); - } else { - chunks.push({ type: "error", message: event.message }); - } - ctx.emit(errorEvent(ctx.conversationId, ctx.turnId, event.message, event.code)); + // Handled by the retry loop in executeStep (not here): an error event + // is intercepted before processEvent so the step can decide whether to + // retry (suppressing the error) or give up (emit it). processEvent + // never receives an "error" event. break; } } @@ -314,34 +320,142 @@ async function executeStep(ctx: StepContext): Promise<StepResult> { // Swallow — D7. } - try { - const opts: ProviderStreamOptions = { - ...ctx.providerOpts, - ...(ctx.turnSpan !== undefined && stepSpan !== undefined ? { logger: stepSpan.log } : {}), - }; - const stream = ctx.provider.stream(ctx.messages, ctx.tools, opts); - for await (const event of stream) { - if (ctx.signal.aborted) break; - processEvent(event, chunks, toolCalls, dispatcher, ctx, stepSpan, timing, toolDispatchTimes); - if (event.type === "usage") { - stepUsage = addUsage(stepUsage, event.usage); + // Retry loop: wrap provider.stream() consumption. Retries are ONLY + // attempted when no content was emitted yet this step (the safety + // invariant — never duplicate partial output). On a retryable error — + // either an EMITTED `error` ProviderEvent with `retryable === true`, OR a + // THROWN error (retryable-by-default when pre-content) — with !hadContent: + // ask retry.delayFor(attempt); if it returns a delay → emit a transient + // provider-retry AgentEvent, sleep via the injected retry.sleep (abortable), + // attempt++, re-call provider.stream(); if it returns undefined (budget + // exhausted) → give up. Non-retryable emitted errors (retryable === false or + // absent), errors after content, and the no-retry-configured case all fall + // through to "give up" — identical to the pre-retry behavior. + let hadContent = false; + let attempt = 0; + while (true) { + let errored = false; + let wasThrown = false; + let errorMessage: string | undefined; + let errorCode: string | undefined; + let errorRetryable: boolean | undefined; + let thrownErr: unknown; + + try { + const opts: ProviderStreamOptions = { + ...ctx.providerOpts, + ...(ctx.turnSpan !== undefined && stepSpan !== undefined ? { logger: stepSpan.log } : {}), + }; + const stream = ctx.provider.stream(ctx.messages, ctx.tools, opts); + for await (const event of stream) { + if (ctx.signal.aborted) break; + if (event.type === "error") { + // Intercept: hold for the retry decision — don't push a chunk + // or emit yet (a successful retry would leave a stale error). + errored = true; + errorMessage = event.message; + errorCode = event.code; + errorRetryable = event.retryable; + break; + } + if ( + event.type === "text-delta" || + event.type === "reasoning-delta" || + event.type === "tool-call" || + event.type === "usage" + ) { + hadContent = true; + } + processEvent( + event, + chunks, + toolCalls, + dispatcher, + ctx, + stepSpan, + timing, + toolDispatchTimes, + ); + if (event.type === "usage") { + stepUsage = addUsage(stepUsage, event.usage); + } + if (event.type === "finish") { + finishReason = event.reason; + } } - if (event.type === "finish") { - finishReason = event.reason; + } catch (err) { + errored = true; + wasThrown = true; + errorMessage = err instanceof Error ? err.message : String(err); + errorCode = undefined; + errorRetryable = undefined; + thrownErr = err; + } + + // Abort (during stream) → stop; the runTurn loop seals aborted. + if (ctx.signal.aborted) { + break; + } + + // No error → step succeeded. + if (!errored) { + break; + } + + // Retryable? A thrown error is retryable-by-default when pre-content; + // an emitted error is retryable ONLY when `retryable === true` (absent + // or false → not retried, per the contract). + const isRetryable = wasThrown ? true : errorRetryable === true; + if (ctx.retry !== undefined && !hadContent && isRetryable) { + const delay = ctx.retry.delayFor(attempt); + if (delay !== undefined) { + // Emit the transient provider-retry event BEFORE the sleep so the + // UI shows "⚠ retrying in Ns…" immediately. Not persisted as a + // chat message — it never pollutes the prompt. + ctx.emit( + providerRetryEvent( + ctx.conversationId, + ctx.turnId, + attempt, + delay, + errorMessage ?? "", + errorCode, + ), + ); + // Abortable sleep. If the signal fires during sleep, the shell's + // sleep rejects — we catch it and break so the turn seals aborted. + try { + await ctx.retry.sleep(delay, ctx.signal); + } catch { + // Abort during sleep (or unexpected sleep failure). + } + if (ctx.signal.aborted) { + break; + } + attempt++; + continue; } + // delayFor returned undefined → budget exhausted → give up. + } + + // Give up: emit the suppressed error and end the step. This is the + // single emission point for a terminal provider error (non-retryable, + // post-content, budget-exhausted, or no-retry-configured). + const message = errorMessage ?? ""; + if (errorCode !== undefined) { + chunks.push({ type: "error", message, code: errorCode }); + } else { + chunks.push({ type: "error", message }); } - } catch (err) { - const message = err instanceof Error ? err.message : String(err); - chunks.push({ type: "error", message }); - ctx.emit(errorEvent(ctx.conversationId, ctx.turnId, message)); + ctx.emit(errorEvent(ctx.conversationId, ctx.turnId, message, errorCode)); finishReason = "error"; - // Close step span with error try { - stepSpan?.end({ err }); + stepSpan?.end({ err: thrownErr ?? new Error(message) }); } catch { // Swallow — D7. } stepSpan = undefined; + break; } // Close timing spans: if no first token was seen, end ttft with firstToken: false @@ -524,6 +638,7 @@ export async function runTurn(input: RunTurnInput): Promise<RunTurnResult> { cwd: input.cwd, now, providerOpts: input.providerOpts, + retry: input.retry, }); totalUsage = addUsage(totalUsage, stepResult.usage); diff --git a/packages/session-orchestrator/src/index.ts b/packages/session-orchestrator/src/index.ts index fa8d9e9..aaafb76 100644 --- a/packages/session-orchestrator/src/index.ts +++ b/packages/session-orchestrator/src/index.ts @@ -12,6 +12,7 @@ export { conversationOpened, conversationStatusChanged, createCompactionService, + createRetryStrategy, createSessionOrchestrator, createWarmService, type EnqueueInput, @@ -34,8 +35,13 @@ export { } from "./orchestrator.js"; export { buildUserMessage, + cumulativeSleepMs, defaultDispatchPolicy, + delayFor, generateTurnId, + RETRY_BUDGET_MS, + RETRY_SCHEDULE_MS, + RETRY_TAIL_MS, resolveReasoningEffort, selectFirstProvider, } from "./pure.js"; diff --git a/packages/session-orchestrator/src/orchestrator.ts b/packages/session-orchestrator/src/orchestrator.ts index a533a16..b4d4b35 100644 --- a/packages/session-orchestrator/src/orchestrator.ts +++ b/packages/session-orchestrator/src/orchestrator.ts @@ -11,6 +11,7 @@ import type { ProviderEvent, ProviderStreamOptions, ReasoningEffort, + RetryStrategy, RunTurnInput, RunTurnResult, ToolContract, @@ -24,6 +25,7 @@ import { createMetricsAccumulator } from "./metrics.js"; import { buildUserMessage, defaultDispatchPolicy, + delayFor, generateTurnId, resolveModelName, resolveReasoningEffort, @@ -325,12 +327,45 @@ export interface SessionOrchestratorBundle { readonly activeConversations: ReadonlySet<string>; } +/** + * The concrete retry strategy wired into every turn's `RunTurnInput.retry`. + * + * `delayFor` is the pure schedule (`5s, 10s, 30s, 60s, 5m, 10m, 15m, 30m`, + * then repeat 30m until 8h cumulative scheduled sleep) — no I/O, no clock. + * `sleep` is the abortable I/O effect: a `setTimeout`-based promise that + * rejects when the turn's abort signal fires (so a retry in flight seals the + * turn `aborted`). The kernel imports no timer; this is the shell-provided I/O. + */ +export function createRetryStrategy(): RetryStrategy { + const sleep = (ms: number, signal: AbortSignal): Promise<void> => { + return new Promise((resolve, reject) => { + if (signal.aborted) { + reject(new Error("aborted")); + return; + } + const timer = setTimeout(() => { + signal.removeEventListener("abort", onAbort); + resolve(); + }, ms); + const onAbort = () => { + clearTimeout(timer); + reject(new Error("aborted")); + }; + signal.addEventListener("abort", onAbort, { once: true }); + }); + }; + return { delayFor, sleep }; +} + export function createSessionOrchestrator( deps: SessionOrchestratorDeps, ): SessionOrchestratorBundle { const activeConversations = new Set<string>(); const subscribers = new Map<string, Set<TurnEventListener>>(); const activeTurns = new Map<string, ActiveTurn>(); + // One stateless retry strategy shared by every turn (delayFor is pure; sleep + // is a stateless setTimeout closure). Wired into each RunTurnInput.retry. + const retryStrategy = createRetryStrategy(); function emitToHub(conversationId: string, event: AgentEvent): void { const turn = activeTurns.get(conversationId); @@ -596,6 +631,7 @@ export function createSessionOrchestrator( turnId, signal: controller.signal, providerOpts, + retry: retryStrategy, ...(turnLogger !== undefined ? { logger: turnLogger } : {}), ...(effectiveCwd !== undefined ? { cwd: effectiveCwd } : {}), ...(deps.now !== undefined ? { now: deps.now } : {}), diff --git a/packages/session-orchestrator/src/pure.test.ts b/packages/session-orchestrator/src/pure.test.ts index 9e5d3c4..2cbe15f 100644 --- a/packages/session-orchestrator/src/pure.test.ts +++ b/packages/session-orchestrator/src/pure.test.ts @@ -2,8 +2,13 @@ import type { ProviderContract } from "@dispatch/kernel"; import { describe, expect, it } from "vitest"; import { buildUserMessage, + cumulativeSleepMs, defaultDispatchPolicy, + delayFor, generateTurnId, + RETRY_BUDGET_MS, + RETRY_SCHEDULE_MS, + RETRY_TAIL_MS, resolveReasoningEffort, selectFirstProvider, } from "./pure.js"; @@ -100,3 +105,59 @@ describe("resolveReasoningEffort", () => { expect(resolveReasoningEffort(undefined, "max")).toBe("max"); }); }); + +describe("retry backoff schedule (delayFor)", () => { + it("emits the stepped head: 5s, 10s, 30s, 60s, 5m, 10m, 15m, 30m", () => { + expect(delayFor(0)).toBe(5_000); + expect(delayFor(1)).toBe(10_000); + expect(delayFor(2)).toBe(30_000); + expect(delayFor(3)).toBe(60_000); + expect(delayFor(4)).toBe(300_000); + expect(delayFor(5)).toBe(600_000); + expect(delayFor(6)).toBe(900_000); + expect(delayFor(7)).toBe(1_800_000); + }); + + it("repeats 30m after the head", () => { + expect(delayFor(8)).toBe(RETRY_TAIL_MS); + expect(delayFor(9)).toBe(RETRY_TAIL_MS); + expect(delayFor(20)).toBe(RETRY_TAIL_MS); + }); + + it("gives up (returns undefined) once cumulative sleep exceeds 8h", () => { + // Head sums to 3,705,000 ms; +1,800,000 per extra step. 8h = 28,800,000. + // attempt 20 cumulative = 3,705,000 + 13*1,800,000 = 27,105,000 (< 8h) → retry. + expect(delayFor(20)).toBe(RETRY_TAIL_MS); + // attempt 21 cumulative = 27,105,000 + 1,800,000 = 28,905,000 (> 8h) → stop. + expect(delayFor(21)).toBeUndefined(); + }); + + it("cumulativeSleepMs matches the sum of the schedule", () => { + expect(cumulativeSleepMs(0)).toBe(5_000); + expect(cumulativeSleepMs(1)).toBe(15_000); + expect(cumulativeSleepMs(7)).toBe(RETRY_SCHEDULE_MS.reduce((a, b) => a + b, 0)); + // 8h budget is 28,800,000 ms. + expect(RETRY_BUDGET_MS).toBe(8 * 60 * 60 * 1000); + // The last retry (attempt 20) keeps cumulative under budget. + expect(cumulativeSleepMs(20)).toBeLessThanOrEqual(RETRY_BUDGET_MS); + // The next (attempt 21) exceeds it. + expect(cumulativeSleepMs(21)).toBeGreaterThan(RETRY_BUDGET_MS); + }); + + it("the full schedule has 21 retries then stops", () => { + const schedule: number[] = []; + let attempt = 0; + while (true) { + const delay = delayFor(attempt); + if (delay === undefined) break; + schedule.push(delay); + attempt++; + } + expect(schedule).toHaveLength(21); + expect(schedule[0]).toBe(5_000); + expect(schedule.at(-1)).toBe(RETRY_TAIL_MS); + // 8 stepped head + 13 tail repeats. + expect(schedule.slice(0, 8)).toEqual([...RETRY_SCHEDULE_MS]); + expect(schedule.slice(8).every((d) => d === RETRY_TAIL_MS)).toBe(true); + }); +}); diff --git a/packages/session-orchestrator/src/pure.ts b/packages/session-orchestrator/src/pure.ts index 9a31e17..a028cbe 100644 --- a/packages/session-orchestrator/src/pure.ts +++ b/packages/session-orchestrator/src/pure.ts @@ -9,6 +9,53 @@ export function buildUserMessage(text: string): ChatMessage { return { role: "user", chunks: [{ type: "text", text }] }; } +// ── Provider-error retry backoff schedule ─────────────────────────────────── +// +// Pure, deterministic delay decision (no I/O, no clock) for retrying retryable +// provider errors (HTTP 429 / 5xx "overloaded"). The concrete `sleep` (I/O) +// is wired in the orchestrator; this owns only the policy. + +/** + * Stepped backoff schedule (ms): 5s, 10s, 30s, 60s, 5m, 10m, 15m, 30m. + * After the head is exhausted, {@link RETRY_TAIL_MS} (30m) repeats. + */ +export const RETRY_SCHEDULE_MS = [ + 5_000, 10_000, 30_000, 60_000, 300_000, 600_000, 900_000, 1_800_000, +] as const; + +/** Tail delay (ms) repeated after the stepped head: 30 minutes. */ +export const RETRY_TAIL_MS = 1_800_000; + +/** Cumulative scheduled-sleep budget (ms) after which retrying gives up: 8h. */ +export const RETRY_BUDGET_MS = 8 * 60 * 60 * 1000; + +/** + * Cumulative scheduled sleep through `attempt` (sum of delay[0..attempt]). + * Pure — no I/O, no clock. + */ +export function cumulativeSleepMs(attempt: number): number { + let sum = 0; + for (let i = 0; i <= attempt; i++) { + sum += i < RETRY_SCHEDULE_MS.length ? (RETRY_SCHEDULE_MS[i] ?? RETRY_TAIL_MS) : RETRY_TAIL_MS; + } + return sum; +} + +/** + * Pure, deterministic delay decision for the retry strategy: given the + * 0-based attempt index, return the delay in ms to sleep before the next + * retry, or `undefined` to stop (cumulative budget exhausted). No I/O, no + * clock — fully testable. Matches the plan's schedule: + * `5s, 10s, 30s, 60s, 5m, 10m, 15m, 30m`, then repeat 30m until 8h of + * cumulative scheduled sleep is reached, then give up. + */ +export function delayFor(attempt: number): number | undefined { + const scheduled = RETRY_SCHEDULE_MS[attempt]; + const delay = scheduled !== undefined ? scheduled : RETRY_TAIL_MS; + if (cumulativeSleepMs(attempt) > RETRY_BUDGET_MS) return undefined; // over budget → stop + return delay; +} + /** * Resolve the reasoning-effort level for a turn: * per-turn override → persisted per-conversation value → default `"high"`. diff --git a/packages/wire/src/index.ts b/packages/wire/src/index.ts index bade977..eecd2f7 100644 --- a/packages/wire/src/index.ts +++ b/packages/wire/src/index.ts @@ -273,6 +273,7 @@ export type AgentEvent = | TurnUsageEvent | TurnStepCompleteEvent | TurnErrorEvent + | TurnProviderRetryEvent | TurnDoneEvent | TurnSealedEvent | TurnSteeringEvent; @@ -429,6 +430,31 @@ export interface TurnErrorEvent { readonly code?: string; } +/** + * A retryable provider error is being retried with backoff. Emitted once per + * scheduled retry, BEFORE the sleep, so the UI can show "⚠ Server overloaded — + * retrying in 5s…" immediately. TRANSIENT: emitted to the frontend but NOT + * persisted into the model's message history (it never pollutes the prompt). + * + * When the retry budget is exhausted, the existing `error` event is emitted and + * the turn seals — so the final failure is still a persisted error. `attempt` is + * 0-based (the Nth retry about to happen); `delayMs` is the scheduled sleep + * before that retry fires. + */ +export interface TurnProviderRetryEvent { + readonly type: "provider-retry"; + readonly conversationId: string; + readonly turnId: string; + /** 0-based: this is the Nth retry about to happen. */ + readonly attempt: number; + /** ms the client should expect to wait before the retry fires. */ + readonly delayMs: number; + /** The endpoint's error verbatim (e.g. "HTTP 429: {…overloaded_error…}"). */ + readonly message: string; + /** The HTTP code when known (e.g. "429"). */ + readonly code?: string; +} + /** The turn has completed (model finished generating). */ export interface TurnDoneEvent { readonly type: "done"; @@ -5,7 +5,43 @@ > Keep this lean and current; do not let it re-accrete a step-by-step changelog. ## Status (current) -`tsc -b` EXIT 0 · biome clean · **1537 vitest** green. +`tsc -b` EXIT 0 · biome clean · **1574 vitest** green. + +## Retry with backoff on retryable provider errors (DONE) +When the upstream LLM API returns a retryable error (HTTP 429 / 5xx "overloaded"), +the kernel now retries `provider.stream()` with a stepped backoff, visibly, until +the 8h cumulative-sleep budget is exhausted — then emits the final error and +seals the turn. Retries fire ONLY when no content was emitted yet this step (the +safety invariant — never duplicate partial output). Plan: +`notes/retry-with-backoff-plan.md`; report: `reports/retry-with-backoff.md`. +- **Architecture (kernel hook + shell policy/I/O):** kernel provides the hook + (`RetryStrategy` contract + the retry loop in `runTurn`); the shell + (session-orchestrator) provides the policy (the schedule) + the I/O (an + abortable `setTimeout` sleep). Kernel imports no timer. `retry?` is optional + → omit = no retry (backward-compatible). +- **New transient `AgentEvent` variant** `provider-retry` (`@dispatch/wire`), + emitted once per scheduled retry BEFORE the sleep so the UI can show + "⚠ retrying in Ns…" immediately; NOT persisted to model history (never + pollutes the prompt). Final failure is still a persisted `error` + seal. +- **Schedule:** `5s,10s,30s,60s,5m,10m,15m,30m`, then repeat 30m until 8h of + cumulative scheduled sleep → ~21 retries then give up. Pure `delayFor(attempt)`. +- **Retry trigger:** emitted `error` with `retryable===true` → retry; + `retryable` false/absent → give up; a THROWN error → retryable-by-default + ONLY when pre-content. All gated on `!hadContent` (text/reasoning/tool-call/usage). +- [x] Verified: `tsc -b` EXIT 0, biome clean, **1574 vitest** pass (+16 new: 11 + kernel retry tests with an injected fake `sleep` + pure `delayFor` + stub + provider — zero `@dispatch/*` mocks; 5 pure schedule tests). Transports + unchanged — transport-ws forwards `AgentEvent` verbatim inside `chat.delta`; + transport-http is generic `JSON.stringify`. Unit-tested only — not yet + live-verified against a real 429. +- **Optional follow-up (roadmap):** the CLI renderer + (`packages/cli/src/render.ts` `renderEvent`) has no `default` case and silently + drops `provider-retry` — the yellow-warning/countdown target is the web + frontend, not the CLI, so non-blocking. Optional: render `provider-retry` in + the CLI as a stderr warning + `delayMs` countdown. +- **Frontend handoff (5d3f, separate repo `../dispatch-web`):** render + `provider-retry` as a yellow warning system-message bubble showing `message` + (+`code`) with the `delayMs` countdown. ## Per-edit LSP diagnostics auto-append (DONE) After a successful `edit_file`, the extension now calls LSP `getDiagnostics` on the |
