mirror of
https://github.com/luckyyzh/pi-agent-integrated.git
synced 2026-10-03 11:09:34 +00:00
feat: integrate Pi backend and Pi Web
This commit is contained in:
@@ -0,0 +1,506 @@
|
||||
# AgentHarness lifecycle
|
||||
|
||||
`AgentHarness` is the orchestration layer above the low-level agent loop. It owns session persistence, runtime configuration, resource resolution, operation locking, and extension-facing mutation semantics.
|
||||
|
||||
This document describes the current direction and implemented behavior. Some extension/session-facade details are planned and called out explicitly.
|
||||
|
||||
## Ultimate lifecycle goal
|
||||
|
||||
Harness listeners and hooks should be able to close over the `AgentHarness` instance and call public harness APIs from any event where those APIs are documented as allowed. Those calls must not corrupt in-flight turn snapshots, reorder persisted transcript entries, lose pending writes, deadlock settlement, or leave the harness in the wrong phase.
|
||||
|
||||
The intended rule is:
|
||||
|
||||
- structural operations remain rejected while busy
|
||||
- queue operations are accepted at documented turn-safe points
|
||||
- runtime config setters update future snapshots without mutating the current provider request
|
||||
- session writes made while busy are durably queued and flushed in deterministic order
|
||||
- getters return latest harness config, not in-flight snapshots
|
||||
- listeners/hooks currently receive no facade; if they close over the raw harness and call settlement APIs such as `waitForIdle()` during the active run, they can deadlock. A future facade should expose `runWhenIdle()` instead.
|
||||
|
||||
`AssistantMessageStream` already decouples provider transport streaming, such as SSE or websocket reads, from downstream event consumption. The harness can therefore await listeners, extension hooks, persistence, and save-point work without blocking the provider transport reader or reintroducing ad hoc event queues. Lifecycle code should prefer explicit awaited sequencing at harness boundaries over fire-and-forget hook/event settlement.
|
||||
|
||||
A final lifecycle hardening pass should prove these guarantees with a broad listener/hook reentrancy test suite.
|
||||
|
||||
## Error handling
|
||||
|
||||
The current split is:
|
||||
|
||||
- low-level capabilities and helpers use `Result<TValue, TError>` where expected failures are contained and must not throw, such as `ExecutionEnv`, filesystem/shell operations, shell-output capture, resource loading, and compaction helpers
|
||||
- high-level mutation/orchestration APIs such as `Session` and `AgentHarness` reject/throw instead of returning bare results that can be ignored
|
||||
- public `AgentHarness` failures are normalized to `AgentHarnessError` where practical; subsystem errors are preserved as `cause`
|
||||
|
||||
Harness events observe committed state. Public mutators validate required input and persistence before committing when practical, then await notifications. If a hook or subscriber fails after commit, the state change is not rolled back and the public method rejects with `AgentHarnessError` code `"hook"`.
|
||||
|
||||
## State model
|
||||
|
||||
The harness separates state into four categories.
|
||||
|
||||
### Harness config
|
||||
|
||||
Harness config is the latest runtime configuration set by the application or extensions:
|
||||
|
||||
- model
|
||||
- thinking level
|
||||
- tools
|
||||
- active tool names
|
||||
- tool context source
|
||||
- resources
|
||||
- stream options
|
||||
- system prompt or system prompt provider
|
||||
|
||||
Getters return harness config. They do not return the snapshot used by an in-flight provider request.
|
||||
|
||||
Setters update harness config immediately, including while a turn is in flight. Changes affect the next turn snapshot, not the currently running provider request.
|
||||
|
||||
`setResources()` accepts concrete resources and emits `resources_update` on every call with shallow-copied current and previous resources. Applications own loading/reloading resources from disk or other sources and should call `setResources()` with new values.
|
||||
|
||||
`getResources()` returns shallow-copied current resources. It is a live config read, not the last turn snapshot.
|
||||
|
||||
### Turn snapshot
|
||||
|
||||
A turn snapshot is the concrete state used for one LLM turn. It is created by `createTurnState()` and contains:
|
||||
|
||||
- persisted session messages
|
||||
- resolved resources
|
||||
- resolved system prompt
|
||||
- model
|
||||
- thinking level
|
||||
- all tools
|
||||
- active tools
|
||||
- resolved tool context
|
||||
- stream options
|
||||
- derived session id
|
||||
|
||||
Static option values are used directly. System-prompt provider callbacks are invoked once per `createTurnState()` call. All logic for that turn uses the same snapshot.
|
||||
|
||||
Resource arrays are shallow-copied when a snapshot is created. Individual skill and prompt-template objects are not deep-copied.
|
||||
|
||||
`toolContext` is application-defined and required when the configured tools require a non-`undefined` context. A static value is reused, while a zero-argument sync or async provider is resolved once for each turn snapshot. Harness tools receive that resolved value when they execute. Individual tools can structurally require only the context fields they use.
|
||||
|
||||
Stream options are shallow-copied when a snapshot is created. `headers` and `metadata` maps are shallow-copied; their values are not deep-copied. Credentials from `getApiKeyAndHeaders()` are resolved per provider request so expiring tokens can refresh, but the configured stream options and derived session id come from the current turn snapshot.
|
||||
|
||||
### Built-in tools
|
||||
|
||||
The package exports `createReadTool()`, `createWriteTool()`, `createEditTool()`, and `createBashTool()`. They perform filesystem and shell operations exclusively through the `ExecutionEnv` supplied in their tool context. Each tool structurally requires the shared `ExecutionToolContext`, containing `env: ExecutionEnv`; applications may provide additional fields. `createReadTool()` accepts an optional image processor for host-provided conversion and resizing without imposing an image-processing dependency on the agent package. `createBashTool()` accepts an async `prepare` hook that can mutate the command, working directory, environment, and environment-inheritance policy using the current tool context.
|
||||
|
||||
### Session
|
||||
|
||||
The session contains persisted entries only. Session reads return persisted state and do not include queued writes.
|
||||
|
||||
`Session.buildContextEntries()` returns the compaction-aware entry sequence used for model context construction. `Session.buildContext()` derives runtime state from the full active branch, then projects those context entries to `AgentMessage[]`. Custom entries are omitted from model context by default; applications can pass `entryProjectors` to the `Session` constructor or `buildContext()` to project selected custom entries into messages. Applications can also pass stacked `entryTransforms`, which run after the default compaction transform, to filter or reorder context entries before projection.
|
||||
|
||||
Session storage implementations must persist leaf changes as `leaf` entries. `setLeafId()` is not an in-memory-only cursor update; it appends a durable entry whose `targetId` is the active tree leaf or `null` for root. Reopening storage must reconstruct the current leaf from the latest persisted leaf-affecting entry.
|
||||
|
||||
### Pending session writes
|
||||
|
||||
Session writes requested while an operation is active are queued as pending session writes. Pending writes are based on session-entry shapes without generated fields (`id`, `parentId`, `timestamp`).
|
||||
|
||||
Pending session writes are always persisted. They are flushed at save points, at operation settlement, and in failure cleanup.
|
||||
|
||||
A public pending-writes/session-facade API is planned but not implemented yet.
|
||||
|
||||
## Operation phases
|
||||
|
||||
The harness has an explicit phase:
|
||||
|
||||
```ts
|
||||
type AgentHarnessPhase = "idle" | "turn" | "compaction" | "branch_summary" | "retry";
|
||||
```
|
||||
|
||||
Structural operations require `phase === "idle"` and synchronously set the phase before the first `await`:
|
||||
|
||||
- `prompt`
|
||||
- `skill`
|
||||
- `promptFromTemplate`
|
||||
- `compact`
|
||||
- `navigateTree`
|
||||
|
||||
Starting another structural operation while the harness is not idle rejects with `AgentHarnessError` code `"busy"`.
|
||||
|
||||
The following operations are allowed during a turn where appropriate:
|
||||
|
||||
- `steer`
|
||||
- `followUp`
|
||||
- `nextTurn`
|
||||
- `abort`
|
||||
- runtime config setters
|
||||
|
||||
Phase/settlement semantics are still provisional and need a full lifecycle pass.
|
||||
|
||||
## Turn execution
|
||||
|
||||
`prompt`, `skill`, and `promptFromTemplate` follow the same flow:
|
||||
|
||||
1. Assert idle and set phase to `"turn"`.
|
||||
2. Create a turn snapshot with `createTurnState()`.
|
||||
3. Derive invocation text from that snapshot.
|
||||
4. Execute the turn with `executeTurn()`.
|
||||
|
||||
`skill` and `promptFromTemplate` resolve their resource from the same snapshot that is passed to the turn. They do not resolve resources separately.
|
||||
|
||||
`steer`, `followUp`, and `nextTurn` accept text plus optional images and create user messages internally. `nextTurn` messages are inserted before the new user message on the next user-initiated turn.
|
||||
|
||||
Queue modes are live, not turn-snapshotted:
|
||||
|
||||
- `getSteeringMode()` / `setSteeringMode()`
|
||||
- `getFollowUpMode()` / `setFollowUpMode()`
|
||||
|
||||
Changing a queue mode during a run affects the next queue drain. Queue drains happen at safe points.
|
||||
|
||||
## Save points
|
||||
|
||||
A save point occurs after an assistant turn and its tool-result messages have completed.
|
||||
|
||||
At a save point the harness:
|
||||
|
||||
1. flushes pending session writes after the agent-emitted messages for that turn
|
||||
2. creates a fresh turn snapshot if the low-level loop may continue
|
||||
3. applies the fresh context/model/thinking-level/stream-options/session-id state before the next provider request
|
||||
|
||||
This lets model, thinking level, tool, resource, stream option, and system prompt changes made during a turn affect the next turn in the same run, while never mutating an in-flight provider request. Because provider transport reading is already decoupled by `AssistantMessageStream`, save-point work and hook settlement can be awaited directly to keep transcript/session ordering deterministic. The loop callbacks are not recreated at save points.
|
||||
|
||||
The low-level loop converts harness `ThinkingLevel` to provider `reasoning` at the provider boundary:
|
||||
|
||||
- `"off"` -> `undefined`
|
||||
- all other thinking levels pass through
|
||||
|
||||
No state refresh is needed on `agent_end` except flushing leftover pending session writes and clearing the operation phase. The exact `settled` event timing is still under review.
|
||||
|
||||
If the system-prompt callback throws while starting `prompt`, `skill`, or `promptFromTemplate`, the operation rejects with `AgentHarnessError` and the harness returns to idle. If it throws from the save-point snapshot created by `prepareNextTurn`, the low-level agent run records an assistant error message.
|
||||
|
||||
## Hooks and events
|
||||
|
||||
The target hook system is described in [hooks.md](./hooks.md).
|
||||
|
||||
Summary:
|
||||
|
||||
- `AgentHarness` emits typed hook events and consumes typed results.
|
||||
- A single hooks implementation owns registration, cleanup, provenance, and result reducers.
|
||||
- Observational and mutation hooks use one event-specific `on()` API; the event result type determines whether a handler may return a result.
|
||||
- Result-producing events are reduced by typed reducer tables; app-specific hooks add reducers only for app-specific result-producing events.
|
||||
- Hook registration provenance is sidecar metadata on the registration. Resource and tool provenance belongs on app-specific concrete value types.
|
||||
- Hook context should be a plain object of facades, not raw internals or late-bound getter mazes.
|
||||
|
||||
Event payloads describe what is happening. Harness getters describe latest config for future snapshots. Hook and listener settlement should be awaited in lifecycle order where possible; transport backpressure is handled below the harness by `AssistantMessageStream`, so the harness does not need a separate async event queue merely to keep SSE or websocket reads flowing.
|
||||
|
||||
### Summarization retry events
|
||||
|
||||
When the harness is configured with a retry policy, generated compaction and branch-summary requests emit retry lifecycle events for transient provider errors:
|
||||
|
||||
- `retry_scheduled`: a retry was scheduled. Includes `operation: "compaction" | "branch_summary"`, `attempt`, `maxAttempts`, `delayMs`, and `errorMessage`.
|
||||
- `retry_attempt_start`: the backoff delay completed and the retried summarization request is starting. Includes `operation`.
|
||||
- `retry_finished`: the retry loop finished after success, exhaustion, or abort. Includes `operation`.
|
||||
|
||||
These events are observational and do not accept hook results.
|
||||
|
||||
## Planned session facade
|
||||
|
||||
Extensions should eventually interact with a harness-scoped `HarnessSession` facade rather than the raw session. The facade should wrap the internal session and enforce harness pending-write ordering semantics. Once this exists, hooks and event listeners can receive a context that exposes the full `AgentHarness` plus the session facade without giving direct access to unordered raw session writes.
|
||||
|
||||
Planned read semantics:
|
||||
|
||||
- reads delegate to persisted session state
|
||||
- reads do not include queued pending writes
|
||||
|
||||
Planned write semantics:
|
||||
|
||||
- idle: persist immediately
|
||||
- busy: enqueue as pending session writes
|
||||
|
||||
A planned diagnostics API may expose pending writes explicitly:
|
||||
|
||||
```ts
|
||||
getPendingWrites(): readonly PendingSessionWrite[]
|
||||
```
|
||||
|
||||
Agent-emitted messages are persisted on `message_end` to preserve transcript ordering. Pending extension/session writes flush after those messages at save points.
|
||||
|
||||
## Abort
|
||||
|
||||
Abort is allowed during a turn. It aborts the low-level run and clears steering/follow-up queues.
|
||||
|
||||
Abort does not clear `nextTurn` messages. Messages queued with `nextTurn()` survive abort and are inserted before the user message on the next user-initiated turn.
|
||||
|
||||
Abort does not discard pending session writes. Pending writes flush at the next save point if reached, at `agent_end`, or in operation failure cleanup.
|
||||
|
||||
Abort barrier semantics still need an audit.
|
||||
|
||||
## Compaction and tree navigation
|
||||
|
||||
Compaction and tree navigation are structural session mutations.
|
||||
|
||||
They are allowed only while idle and are not queued. They operate on persisted session state. The next prompt creates a fresh turn snapshot.
|
||||
|
||||
Branch summary generation is part of the tree navigation operation.
|
||||
|
||||
Auto-compaction and retry decision points are not implemented in `AgentHarness` yet.
|
||||
|
||||
## Test organization
|
||||
|
||||
Harness tests should stay focused by area instead of growing one large catch-all file.
|
||||
|
||||
Current structure:
|
||||
|
||||
- `packages/agent/test/harness/agent-harness.test.ts`: core lifecycle and public API behavior.
|
||||
- `packages/agent/test/harness/agent-harness-stream.test.ts`: stream options and provider hook semantics.
|
||||
|
||||
Preferred future structure:
|
||||
|
||||
- `agent-harness-resources.test.ts`: resource snapshot/loading semantics.
|
||||
- `agent-harness-tools.test.ts`: tool registry getters, active-tool semantics, and update events.
|
||||
- `agent-harness-lifecycle.test.ts`: phase/save-point/settled/reentrancy behavior.
|
||||
|
||||
Use the `pi-ai` faux provider (`registerFauxProvider`, `fauxAssistantMessage`) for deterministic harness/provider tests. Faux response factories can inspect `StreamOptions`, invoke `options.onPayload`, and return scripted assistant messages without real provider APIs or network access.
|
||||
|
||||
Harness coverage is configured separately from the default package test run:
|
||||
|
||||
```bash
|
||||
npm run test:harness
|
||||
npm run coverage:harness
|
||||
```
|
||||
|
||||
`coverage:harness` runs `test/harness/**/*.test.ts` and reports coverage for `src/harness/**/*.ts` plus the non-harness runtime files it directly exercises (`src/agent.ts` and `src/agent-loop.ts`) into `coverage/harness`. Type-only dependencies such as `src/types.ts` are not included because they have no meaningful runtime coverage.
|
||||
|
||||
## Implementation todo
|
||||
|
||||
This list tracks the remaining work before treating `AgentHarness` as migration-ready. Active/planned items are ordered from easiest to hardest. Completed items are archived at the bottom.
|
||||
|
||||
### 1. Add explicit tool registry read/update semantics
|
||||
|
||||
Status: In progress
|
||||
|
||||
Done:
|
||||
|
||||
- Added `setTools(tools, activeToolNames?)`.
|
||||
- Added `setActiveTools(toolNames)`.
|
||||
- Invalid active tool names reject with `AgentHarnessError`.
|
||||
- Added generic app tool and context shapes via `AgentHarness<TContext, TSkill, TPromptTemplate, TTool>`.
|
||||
- Exported `QueueMode` from core types.
|
||||
- Added `AgentHarnessOptions.steeringMode` and `followUpMode`.
|
||||
- Added live `getSteeringMode()` / `setSteeringMode()` and `getFollowUpMode()` / `setFollowUpMode()`.
|
||||
- Added `getTools()` and `getActiveTools()`.
|
||||
- Added `tools_update` observability events, including active-tool-only updates.
|
||||
- Active tool changes are persisted as branch-scoped `active_tools_change` entries.
|
||||
- Duplicate tool names and duplicate active tool names reject.
|
||||
|
||||
Remaining:
|
||||
|
||||
- None.
|
||||
|
||||
Notes:
|
||||
|
||||
- Observability design: [observability.md](./observability.md)
|
||||
|
||||
### 2. Design per-`AgentHarness` model registry
|
||||
|
||||
Status: Planned
|
||||
|
||||
Done:
|
||||
|
||||
- Current `setModel()` behavior is preserved.
|
||||
|
||||
Remaining:
|
||||
|
||||
- Decide how applications supply the model registry.
|
||||
- Decide whether the harness stores concrete `Model` objects, model references, or both.
|
||||
- Validate model selection against the registry.
|
||||
- Define model change semantics during active turns and save points.
|
||||
|
||||
### 3. Full `AgentHarness` lifecycle/state pass
|
||||
|
||||
Status: In progress
|
||||
|
||||
Done:
|
||||
|
||||
- Removed constructor `void syncFromTree()`, `syncFromTree()`, `liveOperationId`, and `shell()`.
|
||||
- Added `createTurnState()`, `applyTurnState()`, and `executeTurn()`.
|
||||
- Added explicit `phase` in place of boolean idle state.
|
||||
- Save points refresh context, model, thinking level, stream options, and session snapshot state.
|
||||
- Pending session writes use session-entry shapes without generated fields.
|
||||
- Pending session writes flush at save points, settlement, and failure cleanup.
|
||||
- `steer`, `followUp`, and `nextTurn` create user messages from text plus optional images.
|
||||
- `nextTurn` messages are inserted before the new user prompt.
|
||||
- Structural compaction/tree operations restore phase with `finally`.
|
||||
- Public harness failures normalize subsystem causes to `AgentHarnessError`.
|
||||
- Pending session writes flush one-by-one and are not dropped on failure.
|
||||
- Queue drains roll back if queue-update notification fails.
|
||||
- `message_end` persistence happens before subscriber notification.
|
||||
- `abort()` signals cancellation before notifications and still waits for idle through notification errors.
|
||||
- Idle model/thinking/tool updates validate and persist before committing in-memory state.
|
||||
- `setLeafId()` persists durable `leaf` entries so tree navigation survives storage reopen.
|
||||
|
||||
Remaining:
|
||||
|
||||
- Finalize phase/idle semantics.
|
||||
- Audit whether `settled` can fire too early.
|
||||
- Make session writes inside `settled` callbacks deterministic.
|
||||
- Audit follow-up behavior around `agent_end`.
|
||||
- Implement auto-compaction decision point.
|
||||
- Implement retry handling.
|
||||
- Verify `before_agent_start` hook semantics against coding-agent.
|
||||
- Decide whether `before_agent_start` needs more turn info such as tools/tool snippets.
|
||||
- Document or change runtime config event timing while busy.
|
||||
- Audit `abort()` barrier semantics.
|
||||
|
||||
### 4. Implement generic hook/event extension mechanism
|
||||
|
||||
Status: Designed in [hooks.md](./hooks.md), not implemented
|
||||
|
||||
Done:
|
||||
|
||||
- Removed `AgentHarnessContext`.
|
||||
- Hooks receive only event payloads.
|
||||
- `emitHook(event)` derives the hook type from `event.type`.
|
||||
- Provider request/payload hooks have ordered transform semantics.
|
||||
|
||||
Remaining:
|
||||
|
||||
- Add `HookEvent`, `ResultOf`, registration options with generic source metadata, and the single `AgentHarnessHooks` implementation.
|
||||
- Move result chaining out of `AgentHarness` into reducer functions.
|
||||
- Type-check base harness reducers so every result-producing `AgentHarnessEvent` has reducer semantics.
|
||||
- Make `AgentHarness` accept and expose the concrete hooks instance with constructor inference for app-specific hooks.
|
||||
- Define the initial harness/context facades exposed through hook context.
|
||||
- Preserve current provider hook behavior, including stream option patch deletion semantics.
|
||||
- Add parity tests for reducer semantics: transform chaining, patch chaining, early block/cancel, cleanup, source metadata, and typed app-specific reducer coverage.
|
||||
|
||||
Notes:
|
||||
|
||||
- Hook design: [hooks.md](./hooks.md)
|
||||
|
||||
### 5. Spike semi-durable harness/session recovery
|
||||
|
||||
Status: Planned
|
||||
|
||||
Done:
|
||||
|
||||
- Wrote durability design: [durable-harness.md](./durable-harness.md)
|
||||
|
||||
Remaining:
|
||||
|
||||
- Decide whether session owns all durable harness state or whether any sidecars are needed for large blobs.
|
||||
- Define durable entries for queues, pending writes, operations, turns, provider requests, and tool calls.
|
||||
- Define resume requirements for app-provided tools, models, extensions, resources, hooks, and auth providers.
|
||||
- Define conservative recovery policy for unfinished agent turns, provider requests, tool calls, compaction, and tree navigation.
|
||||
- Prototype reducer-based recovery from session entries.
|
||||
- Decide whether interrupted operations append user-visible messages or only internal operation entries.
|
||||
|
||||
Notes:
|
||||
|
||||
- Provider streams are not resumable; recovery should restart from durable boundaries or mark operations interrupted.
|
||||
- Unfinished tool calls are unsafe to retry unless tools declare idempotent/retry-safe behavior.
|
||||
|
||||
### 6. Final lifecycle hardening suite
|
||||
|
||||
Status: Planned
|
||||
|
||||
Done:
|
||||
|
||||
- None.
|
||||
|
||||
Remaining:
|
||||
|
||||
- Add broad listener/hook reentrancy tests across relevant events.
|
||||
- Test runtime config setters from low-level lifecycle events and harness events.
|
||||
- Test runtime config observability for model, thinking, resources, tools, active tools, and stream options.
|
||||
- Test resource/tool/model/thinking/stream-option updates during active turns and save points.
|
||||
- Test session writes from listeners and hooks, including `settled` writes.
|
||||
- Test queue operations from turn events, tool events, and provider hooks.
|
||||
- Test rejected structural operations while busy.
|
||||
- Test abort from listeners/hooks.
|
||||
- Test getter behavior during active operations.
|
||||
- Test deterministic ordering of agent-emitted messages and pending listener writes.
|
||||
- Test no deadlocks when async listeners call harness APIs and await them.
|
||||
- Test phase cleanup through success, provider error, hook error, abort, compaction, and tree navigation.
|
||||
|
||||
### 7. Later coding-agent migration plan
|
||||
|
||||
Status: Planned
|
||||
|
||||
Done:
|
||||
|
||||
- None.
|
||||
|
||||
Remaining:
|
||||
|
||||
- Map coding-agent resources to sourced loaders.
|
||||
- Keep app-level resource dedupe/provenance outside the harness.
|
||||
- Adapt extension loading to the future hook/session facade.
|
||||
- Preserve UI/session behavior outside core.
|
||||
- Move coding-agent stream/auth/retry/header behavior onto harness stream configuration and provider hooks.
|
||||
|
||||
---
|
||||
|
||||
## Completed implementation todo
|
||||
|
||||
### 8. Remove `Agent` dependency from `AgentHarness`
|
||||
|
||||
Status: Done
|
||||
|
||||
Done:
|
||||
|
||||
- `AgentHarness` calls `runAgentLoop()` directly.
|
||||
- Harness owns run lifecycle, abort controller, queue draining, provider stream config, event reduction, session persistence, pending write flushing, and save-point snapshots.
|
||||
- Harness tests cover prompt construction, queue draining, abort behavior, save-point refresh, pending write ordering, awaited listener settlement, tool hooks, and provider stream wrapping.
|
||||
|
||||
Remaining:
|
||||
|
||||
- None.
|
||||
|
||||
Notes:
|
||||
|
||||
- Broader listener/hook reentrancy coverage is tracked in item 6.
|
||||
|
||||
### 9. Finish curated provider/stream configuration
|
||||
|
||||
Status: Done
|
||||
|
||||
Done:
|
||||
|
||||
- Added curated `AgentHarnessOptions.streamOptions`, `getStreamOptions()`, and `setStreamOptions()`.
|
||||
- Stream options, headers, metadata, and derived session id are snapshotted per turn.
|
||||
- Harness-owned stream wrapper calls `streamSimple()` and keeps lifecycle-owned `signal` and `reasoning` from the low-level loop.
|
||||
- `getApiKeyAndHeaders()` resolves credentials per provider request.
|
||||
- `before_provider_request`, `before_provider_payload`, and `after_provider_response` hooks are implemented.
|
||||
- Stream option patching supports explicit field deletion and ordered hook chaining.
|
||||
- `agent-harness-stream.test.ts` covers forwarding, auth merge, hook patching/deletion/chaining, payload hooks, and busy/save-point snapshot behavior.
|
||||
|
||||
Remaining:
|
||||
|
||||
- None.
|
||||
|
||||
### 10. Complete low-level `Result` cleanup
|
||||
|
||||
Status: Done
|
||||
|
||||
Done:
|
||||
|
||||
- Added generic `Result<TValue, TError>` plus helpers.
|
||||
- Updated `ExecutionEnv` and `NodeExecutionEnv` to return typed results for filesystem/process operations.
|
||||
- Split filesystem and shell capabilities.
|
||||
- Moved JSONL session storage/repo onto filesystem picks instead of direct Node imports.
|
||||
- Added `ExecutionEnv.appendFile()` for streaming append use cases.
|
||||
- Updated skill and prompt-template loaders to consume `ExecutionEnv` results.
|
||||
- Updated shell output capture to return a result and use `ExecutionEnv`, including full-output spill via `appendFile()`.
|
||||
- Removed `NodeExecutionEnv` from browser-safe root exports.
|
||||
- Replaced `Buffer` usage in generic truncation utilities with runtime-neutral UTF-8 handling.
|
||||
- Converted compaction and branch-summary helpers to typed result returns.
|
||||
- Added `readTextLines()` so JSONL metadata loading reads only the header line.
|
||||
- Removed no-op abort handling from Node filesystem methods where cancellation is not meaningful.
|
||||
- Mapped filesystem errors crossing the session boundary to typed `SessionError`.
|
||||
- Added typed branch-summary errors and cause-aware public harness error normalization.
|
||||
- Resource loaders report structured diagnostics for non-`not_found` filesystem failures.
|
||||
- Expanded `NodeExecutionEnv` tests for file operations, exec errors, aborts, callbacks, timeouts, and shell-output spill.
|
||||
|
||||
Remaining:
|
||||
|
||||
- None.
|
||||
|
||||
Notes:
|
||||
|
||||
- Keep low-level capability/helper APIs non-throwing where they return `Result`.
|
||||
- Keep session storage/repo/session APIs throwing typed `SessionError`.
|
||||
- Keep public structural harness failures normalized to `AgentHarnessError`.
|
||||
- Keep Node-specific APIs isolated under `src/harness/env/nodejs.ts`, Node-backed storage/session implementations, or explicit Node-only entry points.
|
||||
- Audit generic harness utilities for Node globals as APIs are added.
|
||||
- Audit package exports so browser/generic imports do not pull Node-only modules.
|
||||
- Keep expanding `ExecutionEnv` and shell-output contract tests as APIs evolve.
|
||||
@@ -0,0 +1,212 @@
|
||||
# Durable AgentHarness and session design
|
||||
|
||||
<!-- Synced from jot zmnps2zu. Edit this file in-repo going forward. -->
|
||||
|
||||
Durable AgentHarness / session design notes.
|
||||
|
||||
## Framing
|
||||
|
||||
A fully durable `AgentHarness` is not realistic by itself because important dependencies are runtime JS supplied by the host app:
|
||||
|
||||
- tool implementations
|
||||
- model/auth providers
|
||||
- extensions and hook handlers
|
||||
- resource loaders
|
||||
- system-prompt callbacks/modifiers
|
||||
|
||||
Tool registries are runtime dependencies. The harness should persist serializable tool configuration, such as active tool names, but not concrete tool implementations.
|
||||
|
||||
The practical target is a semi-durable harness:
|
||||
|
||||
- session is the durable append-only state tree
|
||||
- harness persists the state it owns into session entries
|
||||
- the host app is responsible for recreating compatible non-persistable dependencies on resume
|
||||
- recovery restarts from durable boundaries, not from an in-flight provider stream
|
||||
|
||||
## Session owns durable state
|
||||
|
||||
Treat session as all durable agent state, not just transcript history.
|
||||
|
||||
Existing session state already includes harness state:
|
||||
|
||||
- model changes
|
||||
- thinking-level changes
|
||||
- active-tool changes
|
||||
- leaf entries
|
||||
- labels
|
||||
- compactions and branch summaries
|
||||
- custom messages and custom entries
|
||||
|
||||
That suggests continuing with one durable session log rather than adding harness sidecars. Sidecars may still be useful for large blobs, but the session entry should remain the source-of-truth reference.
|
||||
|
||||
## What the app must provide on resume
|
||||
|
||||
The app must recreate compatible runtime dependencies:
|
||||
|
||||
- model registry / model objects
|
||||
- tool registry
|
||||
- extension set, versions, and ordering
|
||||
- resource loaders
|
||||
- system prompt providers/hooks
|
||||
- auth providers
|
||||
- app-specific hooks
|
||||
|
||||
Harness can validate stable IDs/versions/hashes when available, but it cannot serialize these dependencies itself.
|
||||
|
||||
## Runtime configuration and restore
|
||||
|
||||
Constructor options remain explicit runtime configuration and do not read session state. Hidden async restore in a constructor would make failure handling ambiguous.
|
||||
|
||||
A future async builder/factory should own durable restore:
|
||||
|
||||
```ts
|
||||
const harness = await AgentHarness.builder()
|
||||
.env(env)
|
||||
.session(session)
|
||||
.model(defaultModel)
|
||||
.tools(runtimeTools)
|
||||
.defaultActiveTools(["read", "edit"])
|
||||
.restore({ missingActiveTools: "fail" });
|
||||
```
|
||||
|
||||
`restore()` should read the active branch, reduce durable harness configuration, apply defaults for missing entries, validate against app-supplied runtime dependencies, construct the harness, and optionally emit `source: "restore"` update events after construction.
|
||||
|
||||
For active tools:
|
||||
|
||||
- `active_tools_change` entries are branch-scoped durable config.
|
||||
- If no `active_tools_change` exists on the branch, restore uses builder defaults, or all registered tools if no default active names were supplied.
|
||||
- Active tool names must be unique.
|
||||
- Tool registry names must be unique.
|
||||
- Missing restored active tool names should fail restore by default; permissive drop/disable policies can be added explicitly later.
|
||||
- Concrete tools are never restored from session; the host app must provide compatible tools.
|
||||
|
||||
## What harness should persist
|
||||
|
||||
Minimum useful durability entries:
|
||||
|
||||
- branch-scoped active tool names
|
||||
- queued steer/followUp/nextTurn messages
|
||||
- queue consumption tied to a turn
|
||||
- pending session writes accepted during active operations
|
||||
- pending write application status
|
||||
- operation start/finish/interruption
|
||||
- turn start/finish
|
||||
- provider request start/finish, if needed for recovery diagnostics
|
||||
- tool call start/finish, if we want safe tool recovery
|
||||
|
||||
Potential entries:
|
||||
|
||||
```ts
|
||||
type DurableHarnessEntry =
|
||||
| QueueEnqueuedEntry
|
||||
| QueueConsumedEntry
|
||||
| PendingWriteEnqueuedEntry
|
||||
| PendingWriteAppliedEntry
|
||||
| OperationStartedEntry
|
||||
| OperationFinishedEntry
|
||||
| OperationInterruptedEntry
|
||||
| TurnStartedEntry
|
||||
| TurnFinishedEntry
|
||||
| ProviderRequestStartedEntry
|
||||
| ProviderRequestFinishedEntry
|
||||
| ToolCallStartedEntry
|
||||
| ToolCallFinishedEntry;
|
||||
```
|
||||
|
||||
Every accepted mutation must be durable before the public API resolves.
|
||||
|
||||
## Recovery model
|
||||
|
||||
On startup:
|
||||
|
||||
1. Host app registers tools/models/extensions/resources/auth/hooks.
|
||||
2. Harness opens session.
|
||||
3. Harness reduces session entries into:
|
||||
- current leaf
|
||||
- conversation branch
|
||||
- harness config, including active tool names
|
||||
- queues
|
||||
- pending writes
|
||||
- active operation/turn/tool state
|
||||
4. Harness validates required runtime dependencies, including restored active tool names against the app-provided tool registry.
|
||||
5. Harness reconciles unfinished operation state.
|
||||
|
||||
Provider streams are not resumable. Recovery can only retry from a durable boundary or mark the operation interrupted.
|
||||
|
||||
## Recovery policies
|
||||
|
||||
Default conservative policy:
|
||||
|
||||
- unfinished agent turn: mark interrupted, preserve durable queues/pending writes, return idle
|
||||
- unfinished provider request: mark interrupted; do not retry automatically
|
||||
- unfinished tool call: append interrupted/error tool result; retry only if the tool declares retry-safe/idempotent
|
||||
- unfinished compaction: rerun if no compaction entry exists
|
||||
- unfinished branch summary/tree navigation: rerun/apply missing summary or leaf entries if safe
|
||||
|
||||
Optional policy:
|
||||
|
||||
```ts
|
||||
recovery: "mark_interrupted" | "retry_unfinished"
|
||||
```
|
||||
|
||||
`retry_unfinished` must be guarded around non-idempotent tool calls.
|
||||
|
||||
## Critical scenarios
|
||||
|
||||
### Queues
|
||||
|
||||
- Crash before `queue_enqueued`: message was not accepted.
|
||||
- Crash after `queue_enqueued`: message is restored.
|
||||
- Crash after queue drain but before durable turn record: risk of loss/duplication.
|
||||
- Required invariant: consumed queue IDs must be recorded in `turn_started` or equivalent before they are considered consumed.
|
||||
|
||||
### Pending writes
|
||||
|
||||
- Crash before `pending_write_enqueued`: write was not accepted.
|
||||
- Crash after enqueue before apply: recovery applies it.
|
||||
- Crash after apply before applied marker: deterministic target entry IDs let recovery detect the entry already exists and mark it applied.
|
||||
|
||||
### Agent loop turn
|
||||
|
||||
- Crash before provider request: retry or mark interrupted.
|
||||
- Crash during provider request: mark interrupted by default.
|
||||
- Crash after provider response before assistant message persisted: response is lost unless provider result was journaled.
|
||||
- Crash after assistant message persisted: recover from durable message.
|
||||
|
||||
### Tool calls
|
||||
|
||||
- Crash after tool call starts but before result: external side effects may already have happened.
|
||||
- Default recovery should not rerun non-idempotent tools.
|
||||
- Tool calls need stable IDs and retry-safety metadata for automatic recovery.
|
||||
|
||||
### Compaction
|
||||
|
||||
- Crash before summary generation: rerun preparation/summary.
|
||||
- Crash after generated summary but before compaction entry: rerun unless summary was journaled.
|
||||
- Crash after compaction entry: operation is complete; append finish marker if missing.
|
||||
|
||||
### Branch summary / tree navigation
|
||||
|
||||
- Crash before summary: rerun or mark interrupted.
|
||||
- Crash after summary entry before leaf entry: append missing leaf entry.
|
||||
- Crash after leaf entry: operation is complete; append finish marker if missing.
|
||||
|
||||
## Minimum viable spike
|
||||
|
||||
1. Add durable queue entries.
|
||||
2. Add durable pending write entries with deterministic target IDs.
|
||||
3. Add operation start/finish/interrupted entries.
|
||||
4. Add turn start with consumed queue IDs.
|
||||
5. Recover by reducing the session log.
|
||||
6. Mark unfinished agent turns interrupted by default.
|
||||
7. Rerun unfinished compaction/tree operations only when no final entry exists.
|
||||
8. Do not retry unfinished tool calls unless tool metadata says retry-safe.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Which remaining harness config entries should move into session first: resources, stream options, system prompt refs?
|
||||
- Should resolved system prompt text be snapshotted per turn for audit/debug?
|
||||
- Do we require strict dependency ID/version matching on resume?
|
||||
- How much provider request data should be journaled?
|
||||
- Should recovery append user-visible assistant interruption messages or only internal operation entries?
|
||||
- Should storage support truncating a final partial JSONL line during recovery?
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,445 @@
|
||||
# AgentHarness hooks design
|
||||
|
||||
<!-- Synced from jot 3utlzkxy. Edit this file in-repo going forward. -->
|
||||
|
||||
Final design.
|
||||
|
||||
## Core model
|
||||
|
||||
Events carry their result type as a type-only phantom:
|
||||
|
||||
```ts
|
||||
declare const HookResult: unique symbol;
|
||||
|
||||
interface HookEvent<TType extends string, TResult = void> {
|
||||
type: TType;
|
||||
readonly [HookResult]?: TResult;
|
||||
}
|
||||
|
||||
type ResultOf<E> = E extends { readonly [HookResult]?: infer R } ? R : void;
|
||||
|
||||
type HookHandler<E, Ctx> = (
|
||||
event: E,
|
||||
ctx: Ctx,
|
||||
signal?: AbortSignal,
|
||||
) => ResultOf<E> | void | Promise<ResultOf<E> | void>;
|
||||
|
||||
type HookObserver<E, Ctx> = (
|
||||
event: E,
|
||||
ctx: Ctx,
|
||||
signal?: AbortSignal,
|
||||
) => void | Promise<void>;
|
||||
```
|
||||
|
||||
Example:
|
||||
|
||||
```ts
|
||||
interface ContextEvent extends HookEvent<"context", { messages?: AgentMessage[] }> {
|
||||
type: "context";
|
||||
messages: AgentMessage[];
|
||||
}
|
||||
|
||||
interface ToolCallEvent extends HookEvent<"tool_call", { block?: boolean; reason?: string }> {
|
||||
type: "tool_call";
|
||||
toolName: string;
|
||||
input: Record<string, unknown>;
|
||||
}
|
||||
|
||||
interface MessageEndEvent extends HookEvent<"message_end"> {
|
||||
type: "message_end";
|
||||
message: AgentMessage;
|
||||
}
|
||||
```
|
||||
|
||||
No result map. No spec table. The event type defines its own result.
|
||||
|
||||
## Hooks interface
|
||||
|
||||
```ts
|
||||
interface AgentHarnessHooks<E extends HookEvent<string, unknown>, Ctx> {
|
||||
context: Ctx;
|
||||
|
||||
setContext(ctx: Ctx): void;
|
||||
|
||||
observe(handler: HookObserver<E, Ctx>): () => void;
|
||||
|
||||
on<TType extends E["type"]>(
|
||||
type: TType,
|
||||
handler: HookHandler<Extract<E, { type: TType }>, Ctx>,
|
||||
): () => void;
|
||||
|
||||
emit<TEvent extends E>(
|
||||
event: TEvent,
|
||||
signal?: AbortSignal,
|
||||
): Promise<ResultOf<TEvent> | undefined>;
|
||||
|
||||
addCleanup(cleanup: () => void | Promise<void>): () => void;
|
||||
|
||||
clear(): Promise<void>;
|
||||
dispose(): Promise<void>;
|
||||
}
|
||||
```
|
||||
|
||||
Important split:
|
||||
|
||||
- `observe()` sees all events, read-only, return ignored.
|
||||
- `on(type, handler)` participates in that event’s semantics.
|
||||
- `emit(event)` is the only thing `AgentHarness` calls.
|
||||
- `clear()` removes observers/handlers and runs cleanups.
|
||||
|
||||
## Default implementation internals
|
||||
|
||||
```ts
|
||||
class DefaultAgentHarnessHooks<E extends HookEvent<string, unknown>, Ctx>
|
||||
implements AgentHarnessHooks<E, Ctx> {
|
||||
context: Ctx;
|
||||
|
||||
private observers = new Set<HookObserver<E, Ctx>>();
|
||||
private handlers = new Map<string, Set<HookHandler<any, Ctx>>>();
|
||||
private cleanups = new Set<() => void | Promise<void>>();
|
||||
|
||||
constructor(ctx: Ctx) {
|
||||
this.context = ctx;
|
||||
}
|
||||
|
||||
setContext(ctx: Ctx): void {
|
||||
this.context = ctx;
|
||||
}
|
||||
|
||||
observe(handler: HookObserver<E, Ctx>): () => void {
|
||||
this.observers.add(handler);
|
||||
return () => this.observers.delete(handler);
|
||||
}
|
||||
|
||||
on(type, handler): () => void {
|
||||
let handlers = this.handlers.get(type);
|
||||
if (!handlers) {
|
||||
handlers = new Set();
|
||||
this.handlers.set(type, handlers);
|
||||
}
|
||||
handlers.add(handler);
|
||||
return () => handlers.delete(handler);
|
||||
}
|
||||
|
||||
async emit(event, signal?) {
|
||||
for (const observer of this.observers) {
|
||||
await observer(event, this.context, signal);
|
||||
}
|
||||
|
||||
switch (event.type) {
|
||||
case "context":
|
||||
return this.emitContext(event, signal);
|
||||
case "before_provider_request":
|
||||
return this.emitBeforeProviderRequest(event, signal);
|
||||
case "before_provider_payload":
|
||||
return this.emitBeforeProviderPayload(event, signal);
|
||||
case "before_agent_start":
|
||||
return this.emitBeforeAgentStart(event, signal);
|
||||
case "tool_call":
|
||||
return this.emitToolCall(event, signal);
|
||||
case "tool_result":
|
||||
return this.emitToolResult(event, signal);
|
||||
case "session_before_compact":
|
||||
case "session_before_tree":
|
||||
return this.emitFirstCancelOrLast(event, signal);
|
||||
default:
|
||||
await this.emitObservationHandlers(event, signal);
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Internal casts are acceptable inside the implementation because `Map<string, ...>` loses specificity. Public API remains typed.
|
||||
|
||||
## Mutation semantics
|
||||
|
||||
### Observation
|
||||
|
||||
```ts
|
||||
await hooks.emit({ type: "message_end", message }, signal);
|
||||
```
|
||||
|
||||
Observers run. `message_end` handlers run. Return ignored unless that event later gets a result type.
|
||||
|
||||
### Context transform
|
||||
|
||||
Handlers run in order. Each sees current messages.
|
||||
|
||||
```ts
|
||||
let current = event;
|
||||
|
||||
for (const handler of handlers("context")) {
|
||||
const result = await handler(current, ctx, signal);
|
||||
if (result?.messages) {
|
||||
current = { ...current, messages: result.messages };
|
||||
}
|
||||
}
|
||||
|
||||
return current.messages === event.messages ? undefined : { messages: current.messages };
|
||||
```
|
||||
|
||||
### Provider request / payload
|
||||
|
||||
Sequential transform. Each handler sees previous output.
|
||||
|
||||
```ts
|
||||
let current = event;
|
||||
|
||||
for (const handler of handlers("before_provider_payload")) {
|
||||
const result = await handler(current, ctx, signal);
|
||||
if (result !== undefined) {
|
||||
current = { ...current, payload: result.payload };
|
||||
}
|
||||
}
|
||||
|
||||
return changed ? { payload: current.payload } : undefined;
|
||||
```
|
||||
|
||||
### Before agent start
|
||||
|
||||
Collect injected messages, chain system prompt.
|
||||
|
||||
```ts
|
||||
let systemPrompt = event.systemPrompt;
|
||||
const messages = [];
|
||||
|
||||
for (const handler of handlers("before_agent_start")) {
|
||||
const result = await handler({ ...event, systemPrompt }, ctx, signal);
|
||||
if (result?.messages) messages.push(...result.messages);
|
||||
if (result?.systemPrompt !== undefined) systemPrompt = result.systemPrompt;
|
||||
}
|
||||
|
||||
return messages.length || systemPrompt !== event.systemPrompt
|
||||
? { messages, systemPrompt }
|
||||
: undefined;
|
||||
```
|
||||
|
||||
### Tool call
|
||||
|
||||
Sequential, early exit on block.
|
||||
|
||||
```ts
|
||||
for (const handler of handlers("tool_call")) {
|
||||
const result = await handler(event, ctx, signal);
|
||||
if (result?.block) return result;
|
||||
}
|
||||
```
|
||||
|
||||
### Tool result
|
||||
|
||||
Sequential patch accumulation. Each handler sees current patched result.
|
||||
|
||||
```ts
|
||||
let current = event;
|
||||
let modified = false;
|
||||
|
||||
for (const handler of handlers("tool_result")) {
|
||||
const result = await handler(current, ctx, signal);
|
||||
if (!result) continue;
|
||||
|
||||
current = {
|
||||
...current,
|
||||
content: result.content ?? current.content,
|
||||
details: result.details ?? current.details,
|
||||
isError: result.isError ?? current.isError,
|
||||
};
|
||||
|
||||
modified = true;
|
||||
}
|
||||
|
||||
return modified
|
||||
? { content: current.content, details: current.details, isError: current.isError }
|
||||
: undefined;
|
||||
```
|
||||
|
||||
### Session-before events
|
||||
|
||||
Sequential, early exit on cancel.
|
||||
|
||||
```ts
|
||||
let last;
|
||||
|
||||
for (const handler of handlers(event.type)) {
|
||||
const result = await handler(event, ctx, signal);
|
||||
if (!result) continue;
|
||||
last = result;
|
||||
if (result.cancel) return result;
|
||||
}
|
||||
|
||||
return last;
|
||||
```
|
||||
|
||||
## Harness usage
|
||||
|
||||
Harness only does this:
|
||||
|
||||
```ts
|
||||
await this.hooks.emit(event, signal);
|
||||
```
|
||||
|
||||
or:
|
||||
|
||||
```ts
|
||||
const result = await this.hooks.emit({ type: "context", messages }, signal);
|
||||
return result?.messages ?? messages;
|
||||
```
|
||||
|
||||
Harness does not store handlers, chain listeners, or know extension policy.
|
||||
|
||||
## Context
|
||||
|
||||
Context is a normal object, not rebuilt per emit.
|
||||
|
||||
```ts
|
||||
const hooks = new CodingAgentHooks({
|
||||
harness: harnessFacade,
|
||||
session: sessionFacade,
|
||||
ui: noUiFacade,
|
||||
});
|
||||
```
|
||||
|
||||
Later:
|
||||
|
||||
```ts
|
||||
hooks.setContext({
|
||||
...hooks.context,
|
||||
ui: tuiFacade,
|
||||
});
|
||||
```
|
||||
|
||||
For dynamic state, prefer stable facades/methods over getter maze:
|
||||
|
||||
```ts
|
||||
interface CodingAgentHookContext {
|
||||
harness: HarnessFacade;
|
||||
session: SessionFacade;
|
||||
ui: UiFacade;
|
||||
models: ModelFacade;
|
||||
}
|
||||
```
|
||||
|
||||
Per-run `signal` is passed as the third handler arg.
|
||||
|
||||
## Extension loading later
|
||||
|
||||
Extension loading can live next to harness and construct hooks:
|
||||
|
||||
```ts
|
||||
const hooks = await loadExtensions({
|
||||
paths,
|
||||
context,
|
||||
hooks: new CodingAgentHooks(context),
|
||||
});
|
||||
const harness = new AgentHarness({ ..., hooks });
|
||||
```
|
||||
|
||||
The loader registers into hooks:
|
||||
|
||||
```ts
|
||||
hooks.on("context", handler);
|
||||
hooks.on("tool_call", handler);
|
||||
hooks.addCleanup(cleanup);
|
||||
```
|
||||
|
||||
For reload:
|
||||
|
||||
```ts
|
||||
await hooks.clear();
|
||||
const nextHooks = await loadExtensions(...);
|
||||
harness.setHooks(nextHooks); // idle-only if supported
|
||||
```
|
||||
|
||||
## Poking holes
|
||||
|
||||
### 1. Error policy must be explicit
|
||||
|
||||
Existing coding-agent catches extension errors, reports them, and continues. New hooks need the same policy, likely:
|
||||
|
||||
```ts
|
||||
errorMode: "continue" | "throw"
|
||||
onError(error)
|
||||
```
|
||||
|
||||
For coding-agent, default should be `"continue"`.
|
||||
|
||||
### 2. Source metadata matters
|
||||
|
||||
Existing runner knows which extension produced an error/resource/tool. Plain `on()` loses that unless we add registration metadata or scopes.
|
||||
|
||||
Probably needed:
|
||||
|
||||
```ts
|
||||
const scope = hooks.createScope({ sourceInfo });
|
||||
scope.on("context", handler);
|
||||
scope.addCleanup(...);
|
||||
```
|
||||
|
||||
Or `on(type, handler, { sourceInfo })`.
|
||||
|
||||
### 3. Some extension capabilities are registries, not hooks
|
||||
|
||||
These are not covered by `emit()` and should stay as registries on `CodingAgentHooks` or an extension host:
|
||||
|
||||
- tools
|
||||
- commands
|
||||
- shortcuts
|
||||
- flags
|
||||
- message renderers
|
||||
- provider registrations
|
||||
- OAuth providers
|
||||
- custom model providers
|
||||
|
||||
That is fine. They do not belong in `AgentHarness`.
|
||||
|
||||
### 4. Existing coding-agent events can be represented
|
||||
|
||||
No blocker for:
|
||||
|
||||
- `context`
|
||||
- `before_provider_request`
|
||||
- `after_provider_response`
|
||||
- `before_agent_start`
|
||||
- `message_end`
|
||||
- `tool_call`
|
||||
- `tool_result`
|
||||
- `input`
|
||||
- `user_bash`
|
||||
- `resources_discover`
|
||||
- `session_before_*`
|
||||
- `session_*`
|
||||
- model/thinking selection events
|
||||
- agent/turn/message/tool lifecycle events
|
||||
|
||||
They become additional event types handled by `CodingAgentHooks`.
|
||||
|
||||
### 5. Need to preserve exact old semantics
|
||||
|
||||
When porting coding-agent, special cases must be copied:
|
||||
|
||||
- `input`: transform chain, `handled` short-circuits.
|
||||
- `user_bash`: first meaningful result wins.
|
||||
- `message_end`: replacement must keep same role.
|
||||
- `before_agent_start`: `ctx.getSystemPrompt()` must reflect current chained prompt.
|
||||
- `resources_discover`: aggregate paths and keep extension source.
|
||||
- `tool_call`: argument mutation remains visible to later handlers.
|
||||
- `tool_result`: later handlers see prior patches.
|
||||
|
||||
The design allows all of that, but the default/coding hooks implementation must encode it.
|
||||
|
||||
### 6. `emit()` switch can miss custom mutation events
|
||||
|
||||
If a subclass adds a result-producing event but forgets to override `emit()`, it will behave observationally. Tests should catch this. Could add a protected strategy registry later if this becomes error-prone, but not initially.
|
||||
|
||||
### 7. Observer semantics are intentionally limited
|
||||
|
||||
Observers see the original emitted event once. They do not see every intermediate mutation. If something needs final transformed state, emit a separate final event or use an event-specific handler.
|
||||
|
||||
## Verdict
|
||||
|
||||
This design can implement a new coding-agent. It is simpler than the current runner, keeps harness clean, and preserves the important extension capabilities as long as `CodingAgentHooks` adds source-aware scopes, registries, cleanup, and the exact old event semantics.
|
||||
|
||||
--- Comments ---
|
||||
|
||||
Thread hn2xk0tzhj on "addCleanup(cleanup"
|
||||
[tmluyaub9v] Owner (2026-05-14T12:55:45.500Z): cleanup should be passed along optionally to on/observe
|
||||
@@ -0,0 +1,964 @@
|
||||
# Models architecture
|
||||
|
||||
This document describes the target design for the next `pi-ai` model/provider refactor. It describes the desired shape, not the current implementation. It is intended to be complete enough to start implementing from a fresh session.
|
||||
|
||||
Goals:
|
||||
|
||||
- `Models` is a dumb runtime collection of providers.
|
||||
- Concrete providers own metadata, auth, model listing, and stream behavior.
|
||||
- API implementations live under `src/api/` and are reusable/lazy.
|
||||
- Concrete provider factories live under `src/providers/`.
|
||||
- Users can import only the providers they need.
|
||||
- Importing a provider must not eagerly import heavy SDKs.
|
||||
- Dynamic model lists are first-class: reads are sync (last-known list), fetching happens in an explicit async `refresh`.
|
||||
- `models.json` and extensions layer by wrapping providers, not by mutating provider internals ad hoc.
|
||||
- Old global APIs survive only in an explicit, temporary `/compat` entrypoint.
|
||||
|
||||
Non-goals for the immediate `pi-ai` pass:
|
||||
|
||||
- Do not migrate coding-agent `ModelRegistry` yet.
|
||||
- Do not keep the stream/API registry inside `Models`.
|
||||
- Do not implement web OAuth flows yet.
|
||||
- Image generation mirrors the chat-side design (`ImagesModels`/`ImagesProvider` in `images-models.ts`); the old global image API (`images.ts`, `images-api-registry.ts`) lives on compat.
|
||||
|
||||
## Package layout
|
||||
|
||||
Target source layout:
|
||||
|
||||
```txt
|
||||
packages/ai/src/
|
||||
index.ts # core exports only; no built-in provider imports
|
||||
models.ts # Models runtime, Provider
|
||||
images-models.ts # ImagesModels runtime, ImagesProvider (mirrors models.ts)
|
||||
compat.ts # temporary old-API compatibility entrypoint
|
||||
auth/ # auth method types, helpers, shared resolveProviderAuth(), login callbacks
|
||||
api/ # API implementations and lazy wrappers
|
||||
openai-completions.ts # real implementation, imports SDKs, exports stream/streamSimple
|
||||
openai-completions.lazy.ts
|
||||
openai-responses.ts
|
||||
openai-responses.lazy.ts
|
||||
openai-codex-responses.ts
|
||||
openai-codex-responses.lazy.ts
|
||||
azure-openai-responses.ts
|
||||
azure-openai-responses.lazy.ts
|
||||
anthropic-messages.ts
|
||||
anthropic-messages.lazy.ts
|
||||
google-generative-ai.ts
|
||||
google-generative-ai.lazy.ts
|
||||
google-vertex.ts
|
||||
google-vertex.lazy.ts
|
||||
mistral-conversations.ts
|
||||
mistral-conversations.lazy.ts
|
||||
bedrock-converse-stream.ts
|
||||
bedrock-converse-stream.lazy.ts
|
||||
openrouter-images.ts # image-generation API implementation
|
||||
openrouter-images.lazy.ts
|
||||
lazy.ts # lazyStream()/lazyApi() helpers
|
||||
(shared helpers: openai-responses-shared, google-shared, transform-messages, ...)
|
||||
providers/ # concrete provider factories and per-provider catalogs
|
||||
openai.ts
|
||||
openai.models.ts # generated OpenAI catalog
|
||||
openai-codex.ts
|
||||
openai-codex.models.ts
|
||||
anthropic.ts
|
||||
anthropic.models.ts
|
||||
google.ts
|
||||
google.models.ts
|
||||
...one pair per built-in provider...
|
||||
openrouter-images.ts # image-generation provider factory
|
||||
faux.ts # test provider factory
|
||||
all.ts # explicit aggregate: builtinModels(), builtinImagesModels(), getBuiltin*()
|
||||
auth/oauth/ # Canonical OAuth implementations (node), lazy-loaded
|
||||
```
|
||||
|
||||
`src/index.ts` must stay core-only. It must not import:
|
||||
|
||||
- generated model catalogs
|
||||
- built-in provider factories
|
||||
- provider SDK implementations
|
||||
- Node-only OAuth modules
|
||||
- `providers/all`
|
||||
- `compat`
|
||||
|
||||
Provider, API, and compat entrypoints are explicit subpath exports.
|
||||
|
||||
## Public usage
|
||||
|
||||
Minimal provider usage:
|
||||
|
||||
```ts
|
||||
import { createModels } from "@earendil-works/pi-ai";
|
||||
import { openaiProvider } from "@earendil-works/pi-ai/providers/openai";
|
||||
|
||||
const models = createModels();
|
||||
models.setProvider(openaiProvider());
|
||||
|
||||
const model = models.getModel("openai", "gpt-4o-mini");
|
||||
if (!model) throw new Error("model not found");
|
||||
|
||||
const response = await models.complete(model, context);
|
||||
```
|
||||
|
||||
Multiple providers:
|
||||
|
||||
```ts
|
||||
const models = createModels();
|
||||
models.setProvider(openaiProvider());
|
||||
models.setProvider(openrouterProvider());
|
||||
```
|
||||
|
||||
All built-ins, explicitly heavy metadata entrypoint:
|
||||
|
||||
```ts
|
||||
import { builtinModels } from "@earendil-works/pi-ai/providers/all";
|
||||
|
||||
const models = builtinModels();
|
||||
```
|
||||
|
||||
`providers/all` may import all provider metadata/catalogs. It still must not eagerly import SDK implementations; provider streams use lazy wrappers.
|
||||
|
||||
## Core runtime: Models
|
||||
|
||||
`Models` is a provider collection plus auth application and stream convenience. No stream registry, no auth resolver strategy object.
|
||||
|
||||
```ts
|
||||
export function createModels(options?: {
|
||||
/** App-owned credential storage. Default: in-memory store. */
|
||||
credentials?: CredentialStore;
|
||||
/** Environment access for auth resolution (env vars, file existence). Default: process.env/node:fs backed; injectable for tests and non-Node hosts. */
|
||||
authContext?: AuthContext;
|
||||
}): MutableModels;
|
||||
|
||||
export interface Models {
|
||||
getProviders(): readonly Provider[];
|
||||
getProvider(id: string): Provider | undefined;
|
||||
|
||||
/** Sync read of last-known models. Best-effort: a provider whose getModels() throws yields no models. */
|
||||
getModels(provider?: string): readonly Model<Api>[];
|
||||
/** Dynamic lists are honestly Model<Api>; narrow with the hasApi() guard. */
|
||||
getModel(provider: string, id: string): Model<Api> | undefined;
|
||||
|
||||
/**
|
||||
* Ask dynamic providers to re-fetch their model lists. With a provider id,
|
||||
* rejects on that provider's failure; without, refreshes all concurrently
|
||||
* best-effort. Static providers are no-ops.
|
||||
*/
|
||||
refresh(provider?: string): Promise<void>;
|
||||
|
||||
/**
|
||||
* Resolve request auth for a model. Includes source label for status UI.
|
||||
* Resolves undefined when the provider is unknown or unconfigured. Rejects
|
||||
* with ModelsError ("oauth" on refresh failure, "auth" on api-key/store
|
||||
* failure); status/availability UIs catch rejections and render
|
||||
* "needs re-login" instead of treating them as unconfigured.
|
||||
*/
|
||||
getAuth(model: Model<Api>): Promise<AuthResult | undefined>;
|
||||
|
||||
stream<TApi extends Api>(
|
||||
model: Model<TApi>,
|
||||
context: Context,
|
||||
options?: ApiStreamOptions<TApi>,
|
||||
): AssistantMessageEventStream;
|
||||
|
||||
complete<TApi extends Api>(
|
||||
model: Model<TApi>,
|
||||
context: Context,
|
||||
options?: ApiStreamOptions<TApi>,
|
||||
): Promise<AssistantMessage>;
|
||||
|
||||
streamSimple(model: Model<Api>, context: Context, options?: SimpleStreamOptions): AssistantMessageEventStream;
|
||||
completeSimple(model: Model<Api>, context: Context, options?: SimpleStreamOptions): Promise<AssistantMessage>;
|
||||
}
|
||||
|
||||
export interface MutableModels extends Models {
|
||||
/** Upsert/replace by provider.id. Provider ids are unique. */
|
||||
setProvider(provider: Provider): void;
|
||||
deleteProvider(id: string): void;
|
||||
clearProviders(): void;
|
||||
}
|
||||
```
|
||||
|
||||
Removed concepts:
|
||||
|
||||
```txt
|
||||
no Models.setStreamFunctions() / getStreamFunctions()
|
||||
no api-registry as a real dispatch mechanism
|
||||
no Models.provider(id) builder, no setModel/upsertModel/patchModel lifecycle
|
||||
no ModelAuthResolver / setAuthResolver — resolution policy is fixed, store is injected
|
||||
```
|
||||
|
||||
If an app needs different auth policy, it wraps providers (wrap auth methods or `getModels`) or passes explicit request auth in stream options.
|
||||
|
||||
## Provider
|
||||
|
||||
A provider is the concrete runtime unit. It owns id/name/base metadata, auth methods, model listing, and stream behavior.
|
||||
|
||||
`Provider` is generic over the APIs its models use. Concrete factories declare what they emit (`openaiProvider(): Provider<"openai-responses" | "openai-completions">`), giving typed model lists to direct factory users. A `Models` collection holds providers as `Provider<Api>`.
|
||||
|
||||
```ts
|
||||
export interface Provider<TApi extends Api = Api> {
|
||||
readonly id: string;
|
||||
readonly name: string;
|
||||
|
||||
readonly baseUrl?: string;
|
||||
readonly headers?: Record<string, string>;
|
||||
|
||||
/**
|
||||
* Required: at least one of apiKey/oauth. Even ambient-credential providers
|
||||
* (env vars, AWS profiles, ADC) and keyless local servers provide apiKey
|
||||
* auth whose resolve() reports whether the provider is configured.
|
||||
* getAuth() returning undefined = not configured.
|
||||
*/
|
||||
readonly auth: ProviderAuth;
|
||||
|
||||
/** Current known models, sync. Static providers: the catalog. Dynamic providers: as of the last refresh (empty before the first). */
|
||||
getModels(): readonly Model<TApi>[];
|
||||
|
||||
/** Dynamic providers only: fetch and update the model list. Concurrent calls share one in-flight fetch. */
|
||||
refreshModels?(): Promise<void>;
|
||||
|
||||
stream<T extends TApi>(model: Model<T>, context: Context, options?: ApiStreamOptions<T>): AssistantMessageEventStream;
|
||||
|
||||
streamSimple(model: Model<TApi>, context: Context, options?: SimpleStreamOptions): AssistantMessageEventStream;
|
||||
}
|
||||
```
|
||||
|
||||
There is no `Provider.api` field. `model.api` carries API identity; the provider dispatches internally (see `createProvider()`).
|
||||
|
||||
`Model.api` remains: existing metadata and tests use it, it is useful for diagnostics, and provider construction uses it for API implementation selection. But `Models` never dispatches on it; the provider does.
|
||||
|
||||
### Typed stream options
|
||||
|
||||
Full stream options are API-specific. `Model<TApi>` pays off by deriving the option type from the API:
|
||||
|
||||
```ts
|
||||
// types.ts — type-only imports from API impl modules are erased, so this is tree-shake safe
|
||||
export interface ApiOptionsMap {
|
||||
"anthropic-messages": AnthropicOptions;
|
||||
"openai-completions": OpenAICompletionsOptions;
|
||||
"openai-responses": OpenAIResponsesOptions;
|
||||
"openai-codex-responses": OpenAICodexResponsesOptions;
|
||||
"azure-openai-responses": AzureOpenAIResponsesOptions;
|
||||
"google-generative-ai": GoogleOptions;
|
||||
"google-vertex": GoogleVertexOptions;
|
||||
"mistral-conversations": MistralOptions;
|
||||
"bedrock-converse-stream": BedrockOptions;
|
||||
}
|
||||
|
||||
export type ApiStreamOptions<TApi extends Api> = TApi extends keyof ApiOptionsMap
|
||||
? ApiOptionsMap[TApi]
|
||||
: StreamOptions & Record<string, unknown>;
|
||||
```
|
||||
|
||||
Custom api strings fall back to the generic shape.
|
||||
|
||||
### Typed model narrowing
|
||||
|
||||
Runtime model lists are dynamic, so `models.getModel()`/`getModels()` honestly return `Model<Api>`. Typing improves at three points:
|
||||
|
||||
1. **`hasApi()` type guard** — runtime-checked narrowing for dynamic lookups (no blind casts):
|
||||
|
||||
```ts
|
||||
export function hasApi<TApi extends Api>(model: Model<Api>, api: TApi): model is Model<TApi>;
|
||||
|
||||
const model = models.getModel("anthropic", "claude-opus-4-7");
|
||||
if (model && hasApi(model, "anthropic-messages")) {
|
||||
// model: Model<"anthropic-messages">, stream options fully typed
|
||||
}
|
||||
```
|
||||
|
||||
2. **`getBuiltinModel()`** — sync, generated-catalog lookup with typed overloads: `(provider, id) -> Model<exact-api-literal>`. The path for hardcoded known models.
|
||||
|
||||
3. **`Provider<TApi>` factories** — typed model lists when using a provider directly, without a `Models` collection.
|
||||
|
||||
Deliberately not done: tying `models.getModel(provider, ...)` to typed provider/model ids would require statically knowing which providers are installed in a mutable runtime collection. The harness path (`streamSimple` + `SimpleStreamOptions`) is API-agnostic and unaffected.
|
||||
|
||||
For comparison: Vercel AI SDK attaches the implementation to the model object, which dissolves dispatch typing but makes models non-serializable (no sessions/RPC/catalogs as plain data), and its `providerOptions` bag is `Record<string, JSON>` checked only by `satisfies` convention. Plain-data models + provider-owned behavior keeps stronger typing where it matters.
|
||||
|
||||
### Name collision
|
||||
|
||||
`types.ts` currently exports `type Provider = KnownProvider | string` (a provider id). Rename that alias to `ProviderId` and fix call sites. The `Provider` interface above takes the name.
|
||||
|
||||
## Provider model listing
|
||||
|
||||
Reads are sync; fetching is an explicit async verb. `Provider.getModels()` returns the current known list — the full catalog for static providers, the last-refreshed list for dynamic ones (llama.cpp, OpenRouter live listing). `refreshModels()` is where dynamic providers fetch.
|
||||
|
||||
This split exists because a sync-or-async union (`Promise<T> | T`) invites latent sync assumptions that detonate on the first async provider, while async-only reads force every consumer (UI lists, extension `find`/`getAll` surfaces) through Promises for data that is almost always static. Sync reads + explicit refresh keeps the staleness visible and the contract single: `getModels()` = last known, `refresh()` = make it current. A fetched list is stale the moment it returns anyway; naming the refresh point is honest about it.
|
||||
|
||||
Apps own the refresh lifecycle: startup, registry reload, opening a model selector. Freshness-critical lookups are two-step: `await models.refresh("llamacpp"); models.getModel("llamacpp", id)`.
|
||||
|
||||
Dynamic refresh must be side-effect-free discovery:
|
||||
|
||||
```txt
|
||||
OK: fetch /v1/models, enumerate local catalog, refresh cached remote model list
|
||||
Not OK: load model, download model, mutate server state, run request probe
|
||||
```
|
||||
|
||||
Provider-specific model lifecycle (load/unload) belongs in app/provider-management commands, not in `refreshModels()`.
|
||||
|
||||
## Streaming path
|
||||
|
||||
`Models.stream()` finds the provider by `model.provider`, resolves auth, merges it into request options, and delegates:
|
||||
|
||||
```ts
|
||||
function stream(model, context, options) {
|
||||
const provider = this.getProvider(model.provider);
|
||||
if (!provider) {
|
||||
// produce an error stream, not a throw — see Error behavior
|
||||
}
|
||||
|
||||
// async setup happens inside the returned stream (lazyStream pattern)
|
||||
const resolution = await this.getAuth(model);
|
||||
const requestModel = resolution?.auth.baseUrl ? { ...model, baseUrl: resolution.auth.baseUrl } : model;
|
||||
const requestOptions = mergeAuth(options, resolution?.auth); // explicit options win per-field
|
||||
|
||||
return provider.stream(requestModel, context, requestOptions);
|
||||
}
|
||||
```
|
||||
|
||||
`stream()` returns `AssistantMessageEventStream` synchronously; async setup (auth resolution, lazy module load) happens inside the returned stream. The forwarding pattern already exists in today's `register-builtins.ts` (`createLazyStream`); extract it as `lazyStream()` in `src/api/lazy.ts`.
|
||||
|
||||
No request hot-path model canonicalization: `stream()` uses the supplied model object as-is. If an app wants fresh model metadata, it refreshes the provider and re-reads (`await models.refresh(p); models.getModel(p, id)`) before starting the turn.
|
||||
|
||||
## API implementations under `src/api`
|
||||
|
||||
An API implementation is reusable stream behavior. It is not a provider.
|
||||
|
||||
Uniform export contract — every real implementation module exports exactly:
|
||||
|
||||
```ts
|
||||
// src/api/anthropic-messages.ts — imports SDKs
|
||||
export function stream(model, context, options) { ... }
|
||||
export function streamSimple(model, context, options) { ... }
|
||||
```
|
||||
|
||||
This makes the module itself satisfy `ProviderStreams`, so the lazy wrapper is one generic helper instead of bespoke per-API plumbing. `ProviderStreams` is the untyped dispatch shape (implementation modules export concretely typed functions, which would not be assignable to a generic method); per-API option typing lives on the modules themselves and on `Provider.stream()` via `ApiStreamOptions`:
|
||||
|
||||
```ts
|
||||
export interface ProviderStreams {
|
||||
stream(model: Model<Api>, context: Context, options?: StreamOptions): AssistantMessageEventStream;
|
||||
streamSimple(model: Model<Api>, context: Context, options?: SimpleStreamOptions): AssistantMessageEventStream;
|
||||
}
|
||||
|
||||
// src/api/lazy.ts
|
||||
export function lazyApi(load: () => Promise<ProviderStreams>): ProviderStreams;
|
||||
|
||||
// src/api/anthropic-messages.lazy.ts
|
||||
export const anthropicMessagesApi = (): ProviderStreams => lazyApi(() => import("./anthropic-messages.ts"));
|
||||
```
|
||||
|
||||
Import chain:
|
||||
|
||||
```txt
|
||||
provider module -> lazy API wrapper -> dynamic import(real API impl) -> SDK deps
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- Bedrock keeps the node-only dynamic import trick (`importNodeOnlyProvider`, `.ts`/`.js` specifier rewrite) inside its lazy wrapper. `setBedrockProviderModule()` (used by the Bun build) moves into the bedrock lazy wrapper module.
|
||||
- Shared helper modules (`openai-responses-shared.ts`, `google-shared.ts`, `transform-messages.ts`, prompt-cache, copilot headers) move to `src/api/` alongside the implementations.
|
||||
|
||||
## Shared API implementations across concrete providers
|
||||
|
||||
Many concrete providers share an API implementation (OpenAI-completions: OpenRouter, Groq, Cerebras, xAI, ZAI, ...). They share lazy API objects by reference:
|
||||
|
||||
```ts
|
||||
import { openAICompletionsApi } from "../api/openai-completions.lazy.ts";
|
||||
|
||||
export function openrouterProvider(): Provider {
|
||||
return createProvider({
|
||||
id: "openrouter",
|
||||
name: "OpenRouter",
|
||||
baseUrl: "https://openrouter.ai/api/v1",
|
||||
auth: { apiKey: envApiKeyAuth("OpenRouter API key", ["OPENROUTER_API_KEY"]) },
|
||||
models: OPENROUTER_MODELS,
|
||||
api: openAICompletionsApi(),
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
This copies Vercel AI SDK's useful property: users import concrete providers; shared protocol implementation is internal.
|
||||
|
||||
## Auth
|
||||
|
||||
Request auth output stays small:
|
||||
|
||||
```ts
|
||||
export interface ModelAuth {
|
||||
apiKey?: string;
|
||||
headers?: Record<string, string>;
|
||||
baseUrl?: string;
|
||||
}
|
||||
```
|
||||
|
||||
If a value cannot be expressed as `apiKey`, `headers`, or `baseUrl`, it is provider config, not auth (Vertex project/location, Bedrock region/profile, Azure apiVersion are provider factory options).
|
||||
|
||||
### Provider auth
|
||||
|
||||
`Provider.auth` has exactly two slots; real providers have at most one api-key path and at most one OAuth path, and the slot names carry the UI's oauth-vs-api-key split without a `kind` discriminant or method ids:
|
||||
|
||||
```ts
|
||||
export interface ProviderAuth {
|
||||
apiKey?: ApiKeyAuth; // stored key/provider env + ambient env/files/ADC/IAM
|
||||
oauth?: OAuthAuth; // login flow + refresh
|
||||
}
|
||||
|
||||
export interface ApiKeyAuth {
|
||||
name: string; // "Anthropic API key"
|
||||
|
||||
/** Interactive setup (prompt for key/provider env). Absent = ambient-only (env, ADC, IAM). */
|
||||
login?(interaction: AuthInteraction): Promise<ApiKeyCredential>;
|
||||
|
||||
/**
|
||||
* Resolve auth from the stored credential and/or ambient sources, merging
|
||||
* per field (credential.key ?? env("..."), credential.env?.NAME ?? env("...")).
|
||||
* undefined = not configured.
|
||||
*/
|
||||
resolve(input: {
|
||||
model: Model<Api>;
|
||||
ctx: AuthContext;
|
||||
credential?: ApiKeyCredential;
|
||||
}): Promise<AuthResult | undefined>;
|
||||
}
|
||||
|
||||
export interface OAuthAuth {
|
||||
name: string; // "Anthropic (Claude Pro/Max)"
|
||||
|
||||
login(interaction: AuthInteraction): Promise<OAuthCredential>;
|
||||
|
||||
/** Exchange the refresh token. Network call; throws on failure (invalid_grant etc.). Runs under the store lock. */
|
||||
refresh(credential: OAuthCredential): Promise<OAuthCredential>;
|
||||
|
||||
/** Side-effect-free derivation of request auth from a valid credential. Covers Copilot-style per-credential baseUrl. Async so lazy wrappers can load the implementation. */
|
||||
toAuth(credential: OAuthCredential): Promise<ModelAuth>;
|
||||
}
|
||||
|
||||
export interface AuthResult {
|
||||
auth: ModelAuth;
|
||||
/** Human-readable label for status UI: "ANTHROPIC_API_KEY", "OAuth", "~/.aws/credentials". */
|
||||
source?: string;
|
||||
}
|
||||
|
||||
export interface AuthContext {
|
||||
env(name: string): Promise<string | undefined>;
|
||||
fileExists(path: string): Promise<boolean>; // supports leading ~
|
||||
}
|
||||
```
|
||||
|
||||
The `refresh`/`toAuth` split lets `Models` own the locked refresh pattern without closure gymnastics: refresh produces a credential, while `toAuth` derives request auth from whatever credential ends up stored.
|
||||
|
||||
OAuth implementations use the provider-neutral `AuthInteraction` protocol directly. A callback-server flow issues a `manual_code` prompt racing the server and aborts the prompt when the callback wins, so the UI needs no provider-specific callback or static callback-server flag.
|
||||
|
||||
### Credentials
|
||||
|
||||
One credential per provider, type-tagged — exactly the shape of today's auth.json (`type: "api_key" | "oauth"` per provider id):
|
||||
|
||||
```ts
|
||||
export interface ApiKeyCredential {
|
||||
type: "api_key";
|
||||
key?: string;
|
||||
env?: ProviderEnv; // e.g. Cloudflare account/gateway ids, Azure/Vertex/Bedrock scoped config
|
||||
}
|
||||
|
||||
export interface OAuthCredential extends OAuthCredentials {
|
||||
type: "oauth"; // access, refresh, expires from OAuthCredentials
|
||||
}
|
||||
|
||||
export type Credential = ApiKeyCredential | OAuthCredential;
|
||||
```
|
||||
|
||||
`ApiKeyCredential.env` stores provider-scoped environment/config values alongside or instead of a key. `ApiKeyAuth.resolve()` merges per field: `credential.key ?? env("CLOUDFLARE_API_KEY")`, `credential.env?.CLOUDFLARE_ACCOUNT_ID ?? env("CLOUDFLARE_ACCOUNT_ID")`, etc. The credential discriminator intentionally matches today's `auth.json` (`api_key`) so the file-backed store does not need lossy type translation.
|
||||
|
||||
### Credential store
|
||||
|
||||
The app injects storage; `pi-ai` ships an in-memory default. Keyed by provider id, one credential per provider:
|
||||
|
||||
```ts
|
||||
export interface CredentialStore {
|
||||
/** Read the stored credential, possibly expired. Display/status use; request auth comes from Models.getAuth(). */
|
||||
read(providerId: string): Promise<Credential | undefined>;
|
||||
|
||||
/**
|
||||
* Serialized write — the only write path. fn sees the current credential
|
||||
* because correct writes (refresh, login-during-refresh) depend on it;
|
||||
* return the new credential, or undefined to leave the entry unchanged.
|
||||
* Mutual exclusion per provider id, cross-process too where the backing
|
||||
* store supports it (file lock). Resolves with the post-write credential.
|
||||
*/
|
||||
modify(
|
||||
providerId: string,
|
||||
fn: (current: Credential | undefined) => Promise<Credential | undefined>,
|
||||
): Promise<Credential | undefined>;
|
||||
|
||||
/** Remove (logout). Serialized against modify. */
|
||||
delete(providerId: string): Promise<void>;
|
||||
}
|
||||
```
|
||||
|
||||
There is deliberately no `set`: an unserialized write path invites read-modify-write races (login-during-refresh clobbering a fresh credential, double token refresh). Call sites:
|
||||
|
||||
```ts
|
||||
await store.modify(pid, async () => credential); // login: store this
|
||||
await store.read(pid); // status UI ("logged in via OAuth")
|
||||
await store.delete(pid); // logout
|
||||
// refresh RMW happens inside Models.getAuth
|
||||
```
|
||||
|
||||
Error semantics: `read` resolves `undefined` for missing entries; methods reject only on storage failure, and `Models` wraps such rejections in `ModelsError` code `"auth"`. Best-effort stores that serve an in-memory view and record persistence errors internally (today's AuthStorage behavior) are valid implementations.
|
||||
|
||||
### Resolution policy (fixed)
|
||||
|
||||
`Models.getAuth(model)` is a decision tree, not a loop. A stored credential owns the provider — ambient/env is consulted only when nothing is stored (AuthStorage parity: no silent env fallback after a failed refresh or for an unmatched credential type):
|
||||
|
||||
```ts
|
||||
const stored = await store.read(provider.id);
|
||||
if (stored) {
|
||||
if (stored.type === "oauth" && provider.auth.oauth) {
|
||||
const oauth = provider.auth.oauth;
|
||||
let credential = stored;
|
||||
if (Date.now() >= credential.expires) { // optimistic check, lock-free
|
||||
const post = await store.modify(provider.id, async (current) => {
|
||||
if (current?.type !== "oauth") return undefined; // logged out meanwhile
|
||||
return Date.now() >= current.expires // authoritative check, under lock
|
||||
? oauth.refresh(current) // throws -> ModelsError("oauth")
|
||||
: undefined; // another process/request refreshed
|
||||
});
|
||||
if (post?.type !== "oauth") return undefined;
|
||||
credential = post;
|
||||
}
|
||||
return { auth: await oauth.toAuth(credential), source: "OAuth" };
|
||||
}
|
||||
if (stored.type === "api_key" && provider.auth.apiKey) {
|
||||
return provider.auth.apiKey.resolve({ model, ctx, credential: stored });
|
||||
}
|
||||
return undefined; // stored credential without matching handler blocks ambient
|
||||
}
|
||||
return provider.auth.apiKey?.resolve({ model, ctx, credential: undefined }); // ambient
|
||||
```
|
||||
|
||||
Properties:
|
||||
|
||||
- Double-checked locking, same as today's `refreshOAuthTokenWithLock`: valid tokens cost one `read` and zero locks; expired tokens lock, re-check under the lock, refresh once globally, persist before release.
|
||||
- Explicit request auth (stream options `apiKey`/`headers`) is merged per-field on top in `stream()`, winning over everything.
|
||||
- Refresh failure rejects with `ModelsError("oauth")`; the stored credential is untouched (preserved for retry). Request paths surface this as a stream error with the real cause ("run /login"); status/availability UIs catch the rejection and render "needs re-login" — documented contract on `getAuth`.
|
||||
|
||||
### Replacing AuthStorage
|
||||
|
||||
The end state for coding-agent: AuthStorage is deleted; its capabilities map onto a `CredentialStore` implementation plus composition.
|
||||
|
||||
Today's `getApiKey` priority and its new home:
|
||||
|
||||
| AuthStorage today | New design |
|
||||
|---|---|
|
||||
| runtime override (CLI `--api-key`) | `withRuntimeOverrides(store, overrides)` decorator: `read` returns the override as an `ApiKeyCredential`; never persisted |
|
||||
| stored `api_key` (with `$ENV`/`!command` via `resolveConfigValue`) | stored `ApiKeyCredential`; config-value resolution happens at `read` in coding-agent's adapter/decorator (command execution stays app policy) |
|
||||
| stored `oauth` + locked refresh, undefined on failure | `getAuth` decision tree above; failure rejects with cause instead of silently unconfiguring |
|
||||
| env var (only when nothing stored) | ambient branch of `apiKey.resolve` |
|
||||
| `fallbackResolver` (models.json custom providers) | gone — custom providers carry their own `auth.apiKey` |
|
||||
|
||||
```txt
|
||||
FileCredentialStore ports AuthStorage's lock backend: read = memory snapshot,
|
||||
modify = withLockAsync(re-read, fn, merge-write), delete,
|
||||
internal error recording (drainErrors equivalent)
|
||||
└─ withConfigValues $ENV / !command at read
|
||||
└─ withRuntimeOverrides --api-key
|
||||
└─ createModels({ credentials: store })
|
||||
|
||||
login/logout UI provider.auth.{oauth,apiKey}.login(interaction) + store.modify/delete
|
||||
status UI store.read(pid) + getAuth try/catch ("needs /login" on rejection)
|
||||
getOAuthProviders presence of provider.auth.oauth across registered providers
|
||||
```
|
||||
|
||||
### Login callbacks
|
||||
|
||||
One interface serves api-key and OAuth login:
|
||||
|
||||
```ts
|
||||
export interface AuthInteraction {
|
||||
/** Aborts the whole login flow. Per-prompt cancellation uses AuthPrompt.signal. */
|
||||
signal?: AbortSignal;
|
||||
|
||||
prompt(prompt: AuthPrompt): Promise<string>;
|
||||
notify(event: AuthEvent): void;
|
||||
}
|
||||
|
||||
/** `signal` lets the flow cancel a pending prompt when an out-of-band event resolves the step. */
|
||||
export type AuthPrompt = { signal?: AbortSignal } & (
|
||||
| { type: "text"; message: string; placeholder?: string }
|
||||
| { type: "secret"; message: string; placeholder?: string }
|
||||
| { type: "select"; message: string; options: readonly { id: string; label: string; description?: string }[] }
|
||||
| { type: "manual_code"; message: string; placeholder?: string }
|
||||
);
|
||||
|
||||
export type AuthEvent =
|
||||
| { type: "auth_url"; url: string; instructions?: string }
|
||||
| { type: "device_code"; userCode: string; verificationUri: string; intervalSeconds?: number; expiresInSeconds?: number }
|
||||
| { type: "progress"; message: string };
|
||||
```
|
||||
|
||||
`prompt()` returns the entered/selected string (`select` returns the option id). Flows race a `manual_code` prompt against a callback server by setting `AuthPrompt.signal` and aborting the prompt when the callback wins.
|
||||
|
||||
### OAuth attachment
|
||||
|
||||
Providers that support OAuth always attach it. There is no factory toggle: the flow is lazy-loaded, so advertising OAuth costs nothing until `login()`/`refresh()` actually runs, and a host that never logs in never loads it.
|
||||
|
||||
```ts
|
||||
export function anthropicProvider(): Provider {
|
||||
return createProvider({
|
||||
id: "anthropic",
|
||||
name: "Anthropic",
|
||||
baseUrl: "https://api.anthropic.com/v1",
|
||||
auth: {
|
||||
apiKey: envApiKeyAuth("Anthropic API key", ["ANTHROPIC_API_KEY"]),
|
||||
oauth: lazyOAuth({
|
||||
name: "Anthropic (Claude Pro/Max)",
|
||||
load: () => import("../auth/oauth/anthropic.ts").then((m) => m.anthropicOAuth),
|
||||
}),
|
||||
},
|
||||
models: ANTHROPIC_MODELS,
|
||||
api: anthropicMessagesApi(),
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
`lazyOAuth()` wraps a dynamically imported `OAuthAuth` so provider definitions can advertise OAuth without importing the implementation (`toAuth` is async for exactly this reason):
|
||||
|
||||
```ts
|
||||
export function lazyOAuth(input: {
|
||||
name: string;
|
||||
load: () => Promise<OAuthAuth>;
|
||||
}): OAuthAuth;
|
||||
```
|
||||
|
||||
OAuth must not force Node-only code (`node:http`, `node:crypto`) into browser bundles: the dynamic import inside `lazyOAuth()` uses the same bundler-opaque variable-specifier trick as the bedrock lazy wrapper. Browser hosts never trigger the load (no stored node OAuth credentials, no login flow). If web OAuth lands later (sitegeist proved feasibility: Web Crypto PKCE, auth tab, fetch token exchange, device-code polling), it is just a different `OAuthAuth` implementation — no reserved option values.
|
||||
|
||||
The built-in flows in `src/auth/oauth/` implement `OAuthAuth` and `AuthInteraction` directly while remaining Node-targeted and lazy-loaded. Copilot derives its credential-specific request endpoint through `toAuth().baseUrl`.
|
||||
|
||||
## Provider wrappers and models.json
|
||||
|
||||
`models.json` is a provider wrapper layer. It does not mutate providers in place:
|
||||
|
||||
```ts
|
||||
function withProviderOverrides(base: Provider, overrides: ProviderOverrides): Provider {
|
||||
return {
|
||||
...base,
|
||||
name: overrides.name ?? base.name,
|
||||
baseUrl: overrides.baseUrl ?? base.baseUrl,
|
||||
headers: mergeHeaders(base.headers, overrides.headers),
|
||||
|
||||
getModels: () => applyModelOverrides(base.getModels(), overrides.models),
|
||||
refreshModels: base.refreshModels?.bind(base),
|
||||
|
||||
stream: base.stream,
|
||||
streamSimple: base.streamSimple,
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
This composes with dynamic providers because `getModels()` delegates to the base source and `refreshModels()` passes through.
|
||||
|
||||
Request-auth config from models.json (`$ENV`, `!command`, inline keys) remains app-owned sidecar state, surfaced either as explicit request auth or as a custom `ApiKeyAuth` the app sets on the wrapped provider's `auth.apiKey`.
|
||||
|
||||
## Custom providers: createProvider()
|
||||
|
||||
One helper builds providers from parts; it handles both single-API and mixed-API providers:
|
||||
|
||||
```ts
|
||||
export function createProvider(input: {
|
||||
id: string;
|
||||
name?: string; // default: id
|
||||
baseUrl?: string;
|
||||
headers?: Record<string, string>;
|
||||
auth: ProviderAuth; // required, at least one of apiKey/oauth (no "no-auth" providers)
|
||||
/** Initial model list (empty for purely dynamic providers). */
|
||||
models: readonly Model<Api>[];
|
||||
/** Dynamic providers: fetch the current list; createProvider stores it and dedupes in-flight calls. */
|
||||
refreshModels?: () => Promise<readonly Model<Api>[]>;
|
||||
/** Single implementation, or map keyed by model.api for mixed-API providers. */
|
||||
api: ProviderStreams | Record<string, ProviderStreams>;
|
||||
}): Provider;
|
||||
```
|
||||
|
||||
- Single `api`: all models stream through it.
|
||||
- Map `api`: `stream()`/`streamSimple()` dispatch on `model.api`; unknown api produces a stream error.
|
||||
|
||||
Mixed-API custom providers must be supported (opencode Go/Zen-style providers expose models backed by different APIs under one provider id).
|
||||
|
||||
Built-in provider factories use `createProvider()` internally. models.json custom providers map onto it directly:
|
||||
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
"my-openai-proxy": {
|
||||
"api": "openai-completions",
|
||||
"baseUrl": "https://proxy.example/v1",
|
||||
"models": [ ... ]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Compat entrypoint
|
||||
|
||||
`@earendil-works/pi-ai/compat` preserves the old global API surface until the coding-agent migration deletes it. New code never imports it.
|
||||
|
||||
Old semantics being preserved: global `stream()` can still dispatch by `model.api` through the legacy api-registry for custom providers, mutated models, and tests/extensions that override a built-in API implementation.
|
||||
|
||||
- `stream/complete/streamSimple/completeSimple(model, ctx, opts)`: real built-in provider/model/api matches route through a singleton `builtinModels()` collection, so provider auth/env/baseUrl behavior is shared with the new runtime. Unknown providers, mutated models, or overridden API registrations fall back to api-registry dispatch plus `getEnvApiKey` injection.
|
||||
- The builtin api registration side effect moves from the root barrel into compat. It skips api ids that already have a registration, since compat may load after a test or extension has already registered an override. `registerApiProvider()/unregisterApiProviders()` keep feeding the compat-local registry; `resetApiProviders()` clears and re-registers builtins.
|
||||
- Sync `getModel/getModels/getProviders` are deprecated aliases of `getBuiltinModel/getBuiltinModels/getBuiltinProviders` from `providers/all` (they were always pure generated-catalog reads — verified: nothing ever mutated the old `modelRegistry`).
|
||||
- Re-exports the per-API lazy stream wrappers (incl. `setBedrockProviderModule`), `env-api-keys.ts`, and the image-generation registry/catalogs; none of these stay on the root barrel.
|
||||
- `export * from "./index.ts"`: compat is a strict superset of the core entrypoint, so consumers switch a file's import path wholesale without symbol surgery.
|
||||
|
||||
coding-agent (and the interim agent package) switch imports of these symbols from `@earendil-works/pi-ai` to `@earendil-works/pi-ai/compat` (import-path-only change) and are otherwise untouched until the ModelManager migration.
|
||||
|
||||
Extension grace period: the coding-agent extension loader (jiti aliases + Bun `virtualModules`) resolves the `@earendil-works/pi-ai` ROOT specifier to the compat entrypoint. Existing user extensions using the old global API (`complete`, `getModel`, `registerApiProvider`, ...) keep working at runtime without changes; they break only when compat is removed at the ModelManager migration, with a migration guide in the changelog. Typechecking is the nudge: editors resolve the root to the slim core types, so extension sources that typecheck must import old globals from `/compat` — which is what the repo example extensions demonstrate.
|
||||
|
||||
## Builtin static helpers
|
||||
|
||||
Typed, sync, generated-catalog-only helpers live with the catalogs (exported from `providers/all`):
|
||||
|
||||
```ts
|
||||
getBuiltinModel(provider, id) // sync, typed overloads from generated catalog
|
||||
getBuiltinModels(provider) // sync
|
||||
getBuiltinProviders() // sync
|
||||
```
|
||||
|
||||
Runtime lookup through a `Models` instance is sync over the last-known provider lists: `models.getModel(...)`. Freshness-critical callers run `await models.refresh(provider)` first.
|
||||
|
||||
Generated catalogs are split per provider (`providers/<id>.models.ts`) by updating `packages/ai/scripts/generate-models.ts`. If the generator change turns out too large for this pass, splitting may be deferred; `providers/all` and provider factories may temporarily import the monolithic `models.generated.ts`, relying on `sideEffects: false` for pruning.
|
||||
|
||||
## Tree-shaking and lazy imports
|
||||
|
||||
Rules:
|
||||
|
||||
1. Main `@earendil-works/pi-ai` import is core-only.
|
||||
2. Provider modules import their catalog, auth helpers, and lazy API wrappers only.
|
||||
3. Lazy API wrappers dynamically import real API implementations.
|
||||
4. Real API implementations import SDK dependencies.
|
||||
5. OAuth implementations are always attached via `lazyOAuth()` and lazy-loaded behind a bundler-opaque dynamic import; provider metadata never eagerly imports Node-only OAuth code.
|
||||
6. `providers/all` imports every built-in provider factory and all catalogs. It is the explicit heavy entrypoint.
|
||||
7. Provider modules are side-effect-free; importing a provider does not register anything globally.
|
||||
8. `package.json` lists only effectful compat/image registration files in `sideEffects`; root and provider modules stay tree-shakeable.
|
||||
9. With code splitting, provider SDKs stay in lazy chunks. Without code splitting, bundlers fold statically reachable lazy API implementations into the single bundle; `providers/all` then pulls all statically visible SDKs. Bedrock is the exception because its AWS SDK implementation is behind a bundler-opaque Node-only import and needs `setBedrockProviderModule()` for standalone single-file bundles.
|
||||
|
||||
Exports map sketch:
|
||||
|
||||
```json
|
||||
{
|
||||
"exports": {
|
||||
".": "./dist/index.js",
|
||||
"./compat": "./dist/compat.js",
|
||||
"./providers/all": "./dist/providers/all.js",
|
||||
"./providers/openai": "./dist/providers/openai.js",
|
||||
"./providers/anthropic": "./dist/providers/anthropic.js",
|
||||
"./providers/*": "./dist/providers/*.js",
|
||||
"./api/*": "./dist/api/*.js"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Browser smoke check (`scripts/check-browser-smoke.mjs`) must keep passing: bundling the core entrypoint (and any non-node provider entrypoint) must not pull `node:http`/`node:crypto`.
|
||||
|
||||
## AgentHarness integration
|
||||
|
||||
`AgentHarness` receives a `Models` instance.
|
||||
|
||||
- `AgentHarnessOptions.models` is required.
|
||||
- The harness does not snapshot `Models` into turn state.
|
||||
- Request path calls `this.models.streamSimple(model, context, options)`; same for compaction/branch-summarization paths.
|
||||
- Request path never calls async `models.getModel()` to canonicalize; if model metadata needs refresh, the app updates the selected model before starting a turn.
|
||||
- Harness tests build `createModels()` and install the faux provider (`fauxProvider()` factory from `providers/faux`).
|
||||
|
||||
## coding-agent next phase (not this pass)
|
||||
|
||||
coding-agent builds providers in layers and binds them per session:
|
||||
|
||||
```txt
|
||||
built-in providers (builtinModels)
|
||||
-> models.json provider wrappers / custom providers (createProvider)
|
||||
-> extension provider wrappers/additions
|
||||
```
|
||||
|
||||
```ts
|
||||
sessionModels.clearProviders();
|
||||
for (const provider of layeredProviders) sessionModels.setProvider(provider);
|
||||
```
|
||||
|
||||
coding-agent owns: `FileCredentialStore` + decorators replacing AuthStorage (see "Replacing AuthStorage"), models.json auth sidecar (`$ENV`, `!command`), command execution policy, provider status labels (from `AuthResult.source`), login/logout UI (driving `auth.{apiKey,oauth}.login()` with `prompt()/notify()`), extension lifecycle, provider-management slash commands.
|
||||
|
||||
Current interim state:
|
||||
|
||||
- `AgentHarness` already accepts a `Models` instance and uses it for turn streaming, compaction, and branch summaries.
|
||||
- coding-agent does not use `AgentHarness` yet; `AgentSession` still drives the low-level `Agent` with a `streamFn`.
|
||||
- coding-agent still uses legacy `AuthStorage` + `ModelRegistry` and imports old global pi-ai APIs through `@earendil-works/pi-ai/compat`.
|
||||
- The extension loader still aliases the pi-ai root to `/compat` as the runtime grace period for old extensions.
|
||||
|
||||
## Implementation TODOs
|
||||
|
||||
Check items off as they land. Keep this list current; it is the working state for resumed sessions.
|
||||
|
||||
### Phase 1 — core types/runtime
|
||||
|
||||
- [x] Rename `types.ts` `Provider` alias to `ProviderId`; fix call sites.
|
||||
- [x] Add `ApiOptionsMap` and `ApiStreamOptions<TApi>` to `types.ts` (type-only imports).
|
||||
- [x] New `models.ts`: `Provider<TApi>` interface, `hasApi()` guard, `ModelsError` + codes. Auth types live in `src/auth/types.ts` (`ProviderAuth` = `{ apiKey?, oauth? }`, credentials, `CredentialStore` (`read`/`modify`/`delete`, one credential per provider), `AuthResult`, `AuthContext`, `ModelAuth`, login callbacks), in-memory store in `src/auth/credential-store.ts`, default context in `src/auth/context.ts` (browser-safe node:fs trick), `lazyStream()` in `src/api/lazy.ts`.
|
||||
- [x] `Models`/`MutableModels`/`createModels({ credentials?, authContext? })` with provider map, sync `getModel(s)` (per-provider failure isolation), explicit async `refresh(provider?)`, `getAuth` (decision tree, double-checked locked refresh), `stream/complete/streamSimple/completeSimple` with per-field auth merge. Tests: `packages/ai/test/models-runtime.test.ts`.
|
||||
- [x] Keep metadata helpers: `calculateCost`, `getSupportedThinkingLevels`, `clampThinkingLevel`, `modelsAreEqual`.
|
||||
|
||||
### Phase 2 — `src/api/`
|
||||
|
||||
- [x] Move stream implementations from `src/providers/` to `src/api/`, renamed by API id (`anthropic.ts` -> `api/anthropic-messages.ts`, etc.).
|
||||
- [x] Normalize each implementation module to export exactly `stream` and `streamSimple`.
|
||||
- [x] Move shared helpers (`openai-responses-shared`, `google-shared`, `transform-messages`, `openai-prompt-cache`, `github-copilot-headers`, `cloudflare`, `simple-options`) to `src/api/`.
|
||||
- [x] Extract `lazyStream()`/`lazyApi()` into `src/api/lazy.ts`.
|
||||
- [x] Add `*.lazy.ts` wrappers per API; bedrock keeps node-only import trick and `setBedrockProviderModule()`.
|
||||
- [x] Delete `providers/register-builtins.ts`. Interim until Phase 5 compat: builtin api-registry registration lives in `stream.ts`; lazy API wrappers are exported from the root barrel.
|
||||
|
||||
### Phase 3 — provider factories + catalogs
|
||||
|
||||
- [x] Auth helpers in `src/auth/helpers.ts`: `envApiKeyAuth()` (with secret-prompt `login`), `lazyOAuth()`. OAuth flow loads go through `auth/oauth/load.ts` (bundler-opaque dynamic import); the `OAuthAuth` exports it references land in Phase 4.
|
||||
- [x] `createProvider()` in `models.ts` (single + mixed `api` map, dispatch on `model.api`, unknown api -> stream error).
|
||||
- [x] Per-provider factories under `src/providers/` for all built-in catalog providers; OAuth attached via `lazyOAuth()` (anthropic, openai-codex, github-copilot); ambient `ApiKeyAuth` for amazon-bedrock (AWS env/profile) and google-vertex (key or ADC+project+location).
|
||||
- [x] `providers/all.ts`: `builtinProviders()`, `builtinModels()`, `getBuiltinModel/getBuiltinModels/getBuiltinProviders` re-exports.
|
||||
- [x] Faux provider factory (`fauxProvider()` in `providers/faux.ts`) for tests; legacy `registerFauxProvider()` kept until compat dies.
|
||||
- [x] Split generated catalogs per provider via `scripts/generate-models.ts` (`providers/<id>.models.ts`); `models.generated.ts` becomes a generated aggregator.
|
||||
|
||||
### Phase 4 — OAuth adaptation
|
||||
|
||||
- [x] Built-in implementations live under `auth/oauth/` and implement `OAuthAuth` directly through `AuthInteraction.prompt()`/`notify()`. They are private provider implementations loaded lazily by provider factories.
|
||||
- [x] Callback-server flows race a `manual_code` prompt, aborted through `AuthPrompt.signal` once the flow settles. The public `oauth` subpath retains only coding-agent extension compatibility types.
|
||||
|
||||
### Phase 5 — packaging
|
||||
|
||||
- [x] `index.ts` core-only and side-effect free (no catalogs, no provider factories, no api-registry, no env-api-keys, no images, no OAuth, no compat). Typed catalog reads (`getBuiltin*`) implemented in `providers/all.ts`; `models.ts` no longer imports `models.generated.ts`.
|
||||
- [x] `compat.ts`: superset of index + old api-dispatch globals, deprecated `getModel/getModels/getProviders` aliases, lazy api wrappers + `setBedrockProviderModule`, `getEnvApiKey`, images. Registration side effect lives here (skip-if-present).
|
||||
- [x] Subpath exports map (`./compat`, `./providers/*`, `./api/*`); `sideEffects` array listing the effectful modules (`compat`, images registration) instead of `false`.
|
||||
- [x] Browser smoke (entry now imports old globals from `/compat`) + shrinkwrap checks green. Internal old-global imports switched to `/compat` already (42 files in agent/coding-agent/examples; vitest configs alias `/compat` to src; spawn-CLI tests resolve workspace dist, so `packages/ai` + `packages/agent` dists were rebuilt).
|
||||
|
||||
### Phase 6 — AgentHarness
|
||||
|
||||
- [x] `AgentHarnessOptions.models` required (`readonly models` on the harness); the harness stream path uses `models.streamSimple()`. `StreamFn` redefined structurally (no compat type dependency); `Models.streamSimple` satisfies it.
|
||||
- [x] Compaction/branch-summarization take the harness `Models` instance. `getApiKeyAndHeaders` is removed entirely — `Models` is the only auth path; per-request key resolution becomes provider auth on the collection. `compact()`/`generateSummary()`/`generateBranchSummary()` lose their explicit `apiKey`/`headers` parameters.
|
||||
- [x] Harness tests use `createModels()` + `fauxProvider()` with unique per-fake provider ids; no global api-registry state, no unregister bookkeeping.
|
||||
|
||||
### Phase 7 — coding-agent bridge (minimal)
|
||||
|
||||
- [x] Switch old-global imports to `@earendil-works/pi-ai/compat` (landed with Phase 5; compat is a superset so the switch was path-only). Extension loader resolves the pi-ai root to compat as the runtime grace period.
|
||||
- [x] Everything else originally sketched here is gated on coding-agent actually streaming through a `Models` instance — coding-agent's `AgentSession` drives the low-level `Agent` via `streamFn`, not the harness — and moved to Phase 9.
|
||||
|
||||
### Phase 8 — wrap-up
|
||||
|
||||
- [x] Update/add tests; run affected suites (tests landed with each phase; `./test.sh` green throughout).
|
||||
- [x] `packages/ai/CHANGELOG.md`: `### Breaking Changes` with migration guide (compat entrypoint, `Provider` -> `ProviderId`, api module moves) + `### Added` for the new Models/provider/auth API.
|
||||
- [x] `packages/coding-agent/CHANGELOG.md`: `### Changed` entry for extension authors — runtime unaffected (loader resolves the pi-ai root to compat), typecheck nudges to `/compat` or the new API; removal happens later with a migration guide.
|
||||
- [x] `packages/agent/CHANGELOG.md`: `### Breaking Changes` for required `AgentHarnessOptions.models`, compaction signature changes, structural `StreamFn`.
|
||||
- [x] `npm run check` clean.
|
||||
|
||||
### Phase 9 — coding-agent on Models + CredentialStore (in scope)
|
||||
|
||||
coding-agent replaces AuthStorage and ModelRegistry's internals with `FileCredentialStore` + a `MutableModels` collection. AgentSession itself stays (AgentHarness adoption is pi 2.0); only its model/auth substrate swaps. Layering is strictly one-directional:
|
||||
|
||||
```txt
|
||||
FileCredentialStore (auth.json, locked, $ENV/!command resolution) + explicit --api-key overlay
|
||||
↑
|
||||
MutableModels: builtin factories (wrapped per models.json config) + custom providers (models.json ∪ extensions)
|
||||
↑
|
||||
ModelRegistry: compatibility facade — sync last-known reads delegate to the collection; registerProvider/login/logout/status for extensions + UI
|
||||
↑
|
||||
AgentSession / sdk / interactive-mode (stream via models; await only auth/refresh paths)
|
||||
```
|
||||
|
||||
Decisions:
|
||||
|
||||
- `AuthStorage` is deleted as a type — it would otherwise depend on provider auth while provider auth depends on its store (circular). Its surface splits: `get`/`set`/`remove` -> `CredentialStore`; `getApiKey` -> `Models.getAuth`; `login`/`logout`/`getAuthStatus` -> ModelRegistry facade methods over `provider.auth.oauth` + the store.
|
||||
- `FileCredentialStore` is self-contained (path, locking, parse/write, chmod, error buffering) and owns `auth.json` semantics, including `$ENV`/`!command` resolution for stored API-key credentials. Persisted values stay raw; resolution returns copies for auth use.
|
||||
- Runtime `--api-key` overrides are an explicit store overlay (an override reads as an ephemeral stored api-key credential, masking stored OAuth — matches today's priority). Every registered provider is guaranteed an `apiKey` auth slot so overrides apply to OAuth-only providers too.
|
||||
- `ModelRegistry.getAll`/`find`/`getAvailable` stay sync for SDK and extension compatibility, delegating to the collection's last-known sync model lists and fast configured-looking status checks. Dynamic providers update through explicit async `refresh()`, and request auth remains async through `getApiKeyAndHeaders()`/`Models.getAuth()`. Extensions also get the collection itself as the forward API.
|
||||
- models.json keeps FULL feature parity, implemented as provider decoration: builtin factories wrapped so `getModels()` applies provider `baseUrl`/`compat` overlays, `modelOverrides`, and custom-model merges (async-safe); provider `apiKey`/`headers`/`authHeader` configs become that provider's `ApiKeyAuth` (config first, factory auth fallback); parse errors keep `getError()` semantics.
|
||||
- Extension `ProviderConfig` parity: provider-keyed `streamSimple`, legacy extension OAuth callbacks adapted to `OAuthAuth`, and full model replacement per provider. Legacy `registerApiProvider` writes stay compat-local for consumers that call global `complete()`; they die with compat.
|
||||
- Copilot: stored-credential baseUrl applied in the wrapped `getModels()` (extension-visible models stay correct) plus per-request `toAuth().baseUrl`.
|
||||
- Cloudflare: provider-auth substitution (key + `CLOUDFLARE_ACCOUNT_ID`/`CLOUDFLARE_GATEWAY_ID` from credential `env` or ambient `AuthContext.env()` -> `ModelAuth.baseUrl`). Built-in compat calls route through `Models`, so they use the same provider auth path.
|
||||
|
||||
Ordering for new sessions:
|
||||
|
||||
1. [x] pi-ai rework first: `Provider.getModels()` sync + optional `refreshModels()`; `Models.getModels`/`getModel` sync, `Models.refresh(provider?)` async; `createProvider` takes `models` array + optional `refreshModels` fetcher (in-flight dedupe). Reverses Phase 1's async-listing decision — see "Provider model listing" for rationale (sync-or-async unions breed latent sync assumptions; async-only breaks sync consumer surfaces like extension `find`/`getAll`).
|
||||
2. [x] Cloudflare provider auth in pi-ai factories: Workers AI and AI Gateway validate their required account/gateway env/config and return resolved `baseUrl`, provider-scoped env, and header suppression/override metadata from provider auth.
|
||||
3. [ ] Add `FileCredentialStore` in coding-agent.
|
||||
- Implement the pi-ai `CredentialStore` interface as a self-contained `auth.json` store; do not depend on the old `AuthStorageBackend` abstraction, though its lock/retry semantics may be ported.
|
||||
- Preserve the existing file format. `ApiKeyCredential` uses `{ type: "api_key", key?, env? }`, matching today's `auth.json`; do not translate `env` into metadata or rewrite discriminators.
|
||||
- Resolve `$ENV`/`!command` in stored API-key `key` and `env` values out of the box using an injected execution/config environment. `$ENV` lookup should come from that environment, and `!command` should run through the shared shell execution path rather than direct `execSync`.
|
||||
- Persist raw config values; resolved credentials returned for auth use must be copies and must not rewrite `$ENV`/`!command` strings unless a caller explicitly stores new values.
|
||||
- `read(provider)` returns the current credential snapshot and records parse/storage errors for status UI parity.
|
||||
- `modify(provider, fn)` must lock, re-read, run `fn`, merge-write the provider entry, chmod `0600`, and return the post-write credential.
|
||||
- `delete(provider)` must lock and remove only that provider's entry.
|
||||
- Add file-backed and in-memory tests covering lock/RMW behavior, `api_key` reads with config-value resolution, OAuth reads, provider `env` preservation, delete, parse errors, and concurrent refresh-style modifications.
|
||||
4. [ ] Add runtime override overlay for coding-agent policy.
|
||||
- `withRuntimeOverrides(store, overrides)` implements CLI `--api-key`: read returns an ephemeral `{ type: "api_key", key }` for each overridden provider, masking stored OAuth/API credentials without persisting.
|
||||
- Runtime overrides must apply even to OAuth-capable providers; every provider registered in coding-agent must retain or gain an `apiKey` auth slot so the overlay is meaningful.
|
||||
- Tests cover precedence: runtime override > stored credential > models.json config auth > ambient provider env, with stored credential blocking ambient fallback.
|
||||
5. [ ] Build provider decoration helpers for `models.json`.
|
||||
- Start from built-in provider factories, not generated model arrays.
|
||||
- Wrap provider `getModels()` so provider-level `baseUrl`/`headers`/`compat`, per-model `modelOverrides`, and custom model merges apply on every sync read.
|
||||
- Preserve `refreshModels()` passthrough so dynamic providers compose with decorations.
|
||||
- Convert provider `apiKey`/`headers`/`authHeader` models.json config into a wrapped `ApiKeyAuth` that resolves config values first and falls back to the base provider auth.
|
||||
- Custom providers with `models` use `createProvider()` with the appropriate lazy API wrapper or extension-provided stream implementation.
|
||||
- Parse errors must keep current `ModelRegistry.getError()` behavior: built-ins remain available, and the error is visible.
|
||||
6. [ ] Copilot `getModels()` baseUrl wrap.
|
||||
- GitHub Copilot OAuth `toAuth()` already returns per-credential request `baseUrl` for streaming.
|
||||
- Wrap Copilot's provider `getModels()` when an OAuth credential is present so extension/UI-visible model metadata also carries the authenticated account base URL.
|
||||
- Keep API-key/env-token Copilot behavior unchanged.
|
||||
- Add tests for model metadata before login, after OAuth credential, after refresh/baseUrl change, and logout.
|
||||
7. [x] Extension OAuth adapter.
|
||||
- Keep only the legacy callback/credential declarations required by coding-agent `ProviderConfig.oauth`.
|
||||
- `login` maps legacy callbacks/events to `AuthInteraction.prompt()`/`notify()`.
|
||||
- `refreshToken` maps to `refresh`; `getApiKey` maps to `toAuth`.
|
||||
- Preserve the type-only pi-ai `oauth` barrel and extension-loader aliases.
|
||||
8. [ ] Rebuild coding-agent `ModelRegistry` over `MutableModels`.
|
||||
- It owns a `MutableModels` instance built from decorated built-ins + models.json custom providers + extension providers.
|
||||
- `getAll()`, `find()`, and `getAvailable()` remain sync compatibility methods over last-known model lists and fast configured-looking auth status. Do not break the extension-facing `modelRegistry` surface for these reads.
|
||||
- `refresh()` is the explicit async freshness boundary: rebuild provider layers and call `models.refresh()` where needed; no global api-registry reset should be part of the new path except compat-only grace behavior.
|
||||
- `registerProvider()`/`unregisterProvider()` mutate provider layers and rebuild the collection.
|
||||
- Facade auth ops (`login`, `logout`, provider status, available OAuth providers) drive `provider.auth.{apiKey,oauth}` and the `CredentialStore`; no `AuthStorage` type remains.
|
||||
- Legacy `registerApiProvider` writes stay only for `/compat` callers and are removed in Phase 10.
|
||||
9. [ ] Rewire consumers.
|
||||
- `AgentSession` stream function resolves through `ModelRegistry`/`Models`, not `getApiKeyAndHeaders()` + compat globals.
|
||||
- SDK options replace `authStorage` with `credentials?: CredentialStore` or an agent-dir-backed default; update `sdk.md` and examples.
|
||||
- `model-resolver`, `--list-models`, model selector, login/logout/status UI, and provider attribution use sync last-known model reads and await only explicit refresh/auth operations.
|
||||
- CLI `--api-key` populates the runtime override decorator instead of mutating `AuthStorage`.
|
||||
- Keep extension loader root-to-compat alias until Phase 10, but expose the new collection/facade as the forward API.
|
||||
10. [ ] Test migration and real-provider validation.
|
||||
- Unit tests for `FileCredentialStore`, runtime override overlay, provider decoration, extension OAuth adapter, Models-backed ModelRegistry facade, and consumer rewiring.
|
||||
- Regression tests for Cloudflare account/gateway env, Copilot OAuth baseUrl wrapping, runtime `--api-key` precedence, `$ENV`/`!command` resolution, and stored credential blocking ambient fallback.
|
||||
- Update existing tests for sync last-known `ModelRegistry.getAll/find/getAvailable` plus explicit async refresh behavior.
|
||||
- Run targeted non-e2e suites plus tmux validation of login flows against real providers (Anthropic OAuth/API key, OpenAI Codex OAuth, GitHub Copilot OAuth, Cloudflare AI Gateway, Bedrock if credentials are available).
|
||||
|
||||
### Phase 10 — compat deletion (pi 2.0 era, separate)
|
||||
|
||||
- [ ] AgentSession -> AgentHarness; the registry facade dies in favor of harness `Models`.
|
||||
- [ ] Move ALL internal `/compat` imports to the new API: every package's src, all tests, and the example extensions (examples then demonstrate the new API). Nothing inside the repo may import `/compat` at that point.
|
||||
- [ ] Delete `/compat`, `env-api-keys.ts`, the extension-loader root-to-compat alias, and the compat-local legacy API registry. The old OAuth registry/provider interface is already gone; the type-only `oauth` barrel remains for extension compatibility.
|
||||
|
||||
### Deferred / follow-ups
|
||||
|
||||
- [ ] Web OAuth implementations (sitegeist-style) as an alternative `OAuthAuth`.
|
||||
- [x] Images API redesign: `ImagesModels`/`ImagesProvider`/`createImagesProvider` mirror the chat-side design (sync reads, explicit refresh, never-reject generation); auth resolution shared with the chat side via the free-standing `resolveProviderAuth()` in `auth/resolve.ts` (which also owns `ModelsError`; both collections pass their store/context as arguments — no resolver object). `openrouterImagesProvider()` factory + `builtinImagesProviders()`/`builtinImagesModels()` in `providers/all`; impl moved to `api/openrouter-images.ts` with a lazy wrapper. The old global image API (registry + `getImageModel*` + `generateImages`) stays on compat; `ImagesProvider` id alias in types.ts renamed to `ImagesProviderId` (mirror of `Provider` -> `ProviderId`).
|
||||
|
||||
## Error behavior
|
||||
|
||||
`undefined` means not found or not configured. Real failures reject or become stream errors.
|
||||
|
||||
```ts
|
||||
export type ModelsErrorCode =
|
||||
| "model_source" // provider model refresh failed
|
||||
| "model_validation" // model object invalid
|
||||
| "provider" // unknown provider, dispatch failure
|
||||
| "stream" // stream setup failure
|
||||
| "auth" // auth resolution failure
|
||||
| "oauth"; // oauth login/refresh failure
|
||||
```
|
||||
|
||||
- `Models.stream()` produces stream errors (error event + error result) for async setup failures; it does not throw after returning the stream.
|
||||
- `Models.getModels()` is a sync best-effort read: a provider whose `getModels()` throws yields no models. `Models.refresh(provider)` rejects on that provider's fetch failure; `Models.refresh()` (all providers) is concurrent best-effort. Apps that need a concrete listing failure refresh the single provider.
|
||||
- Auth resolution and credential store failures reject loudly (`ModelsError` codes `auth`/`oauth`); silent fallback to a different auth path after a failure risks billing surprises. A stored credential always blocks ambient/env fallback, including after a failed refresh.
|
||||
- Status/availability UIs catch `getAuth` rejections and render "needs re-login"; they do not treat rejection as "unconfigured".
|
||||
@@ -0,0 +1,376 @@
|
||||
<!-- Synced from jot qe0ikdqs. Edit this file in-repo going forward. -->
|
||||
|
||||
# Pi Observability Design Notes
|
||||
|
||||
## Goal
|
||||
|
||||
Make `packages/ai` and `packages/agent`/harness observable without depending on OpenTelemetry, Sentry, or any APM vendor.
|
||||
|
||||
Pi should emit stable, structured lifecycle events. External listeners can convert those events into OTel spans, Sentry spans, logs, metrics, or custom telemetry.
|
||||
|
||||
## Mental model
|
||||
|
||||
A trace is one causal tree of work, e.g. one user turn.
|
||||
|
||||
A span is one timed operation in that tree. It is normally represented by IDs, not object pointers:
|
||||
|
||||
```ts
|
||||
interface SpanRecord {
|
||||
traceId: string;
|
||||
spanId: string;
|
||||
parentSpanId?: string;
|
||||
name: string;
|
||||
startTime: number;
|
||||
endTime?: number;
|
||||
attributes: Record<string, unknown>;
|
||||
status: "ok" | "error";
|
||||
}
|
||||
```
|
||||
|
||||
Example tree:
|
||||
|
||||
```text
|
||||
traceId=t1 spanId=s1 parent=- name=pi.agent.prompt
|
||||
traceId=t1 spanId=s2 parent=s1 name=pi.agent.turn
|
||||
traceId=t1 spanId=s3 parent=s2 name=pi.ai.provider.request
|
||||
traceId=t1 spanId=s4 parent=s2 name=pi.agent.tool_call
|
||||
traceId=t1 spanId=s5 parent=s4 name=pi.session.append_entry
|
||||
```
|
||||
|
||||
## Async context
|
||||
|
||||
JavaScript has one event loop but multiple async chains can interleave. A single global `currentContext` breaks under concurrency.
|
||||
|
||||
`AsyncLocalStorage` is the Node equivalent of `ThreadLocal` for async continuations. It lets concurrent operations keep distinct current contexts:
|
||||
|
||||
```ts
|
||||
await Promise.all([
|
||||
runWithPiContext({ userId: "alice" }, () => harness.prompt("A")),
|
||||
runWithPiContext({ userId: "bob" }, () => harness.prompt("B")),
|
||||
]);
|
||||
```
|
||||
|
||||
Deep code can then read the correct current context for the active async chain.
|
||||
|
||||
Pi must run in Node, Bun, browser, workers, and other JS runtimes, so ALS cannot be the core abstraction. It should be a runtime adapter.
|
||||
|
||||
## Core design
|
||||
|
||||
Pi owns a small runtime-agnostic observability abstraction:
|
||||
|
||||
```ts
|
||||
export interface PiObservabilityContext {
|
||||
traceId?: string;
|
||||
currentSpanId?: string;
|
||||
userContext?: Record<string, unknown>;
|
||||
}
|
||||
|
||||
export interface PiObservabilityEvent {
|
||||
type: "start" | "end" | "error" | "event";
|
||||
name: string;
|
||||
traceId: string;
|
||||
spanId?: string;
|
||||
parentSpanId?: string;
|
||||
timestamp: number;
|
||||
durationMs?: number;
|
||||
context?: Record<string, unknown>;
|
||||
payload?: Record<string, unknown>;
|
||||
error?: { name: string; message: string };
|
||||
}
|
||||
|
||||
export interface PiObservability {
|
||||
getContext(): PiObservabilityContext | undefined;
|
||||
runWithContext<T>(context: PiObservabilityContext, fn: () => T): T;
|
||||
emit(event: PiObservabilityEvent): void;
|
||||
hasSubscribers(): boolean;
|
||||
}
|
||||
```
|
||||
|
||||
Public API:
|
||||
|
||||
```ts
|
||||
export function configurePiObservability(observability: PiObservability): void;
|
||||
export function subscribePiObservability(listener: (event: PiObservabilityEvent) => void): () => void;
|
||||
export function runWithPiContext<T>(userContext: Record<string, unknown>, fn: () => T): T;
|
||||
export function traceOperation<T>(name: string, payload: Record<string, unknown>, fn: () => T): T;
|
||||
```
|
||||
|
||||
`traceOperation()`:
|
||||
|
||||
1. reads the current context
|
||||
2. creates `traceId` if missing
|
||||
3. creates a new `spanId`
|
||||
4. uses current span as `parentSpanId`
|
||||
5. emits `start`
|
||||
6. runs callback under child context
|
||||
7. emits `end` or `error`
|
||||
8. rethrows on error
|
||||
|
||||
Pseudo-code:
|
||||
|
||||
```ts
|
||||
function traceOperation<T>(name: string, payload: Record<string, unknown>, fn: () => T): T {
|
||||
const parent = getContext();
|
||||
const traceId = parent?.traceId ?? createId();
|
||||
const spanId = createId();
|
||||
const parentSpanId = parent?.currentSpanId;
|
||||
|
||||
const child = { ...parent, traceId, currentSpanId: spanId };
|
||||
|
||||
emit({ type: "start", name, traceId, spanId, parentSpanId, timestamp: Date.now(), context: parent?.userContext, payload });
|
||||
|
||||
return runWithContext(child, () => {
|
||||
try {
|
||||
const result = fn();
|
||||
// Promise-aware implementation emits end/error after settlement.
|
||||
emit({ type: "end", name, traceId, spanId, parentSpanId, timestamp: Date.now(), context: child.userContext, payload });
|
||||
return result;
|
||||
} catch (error) {
|
||||
emit({ type: "error", name, traceId, spanId, parentSpanId, timestamp: Date.now(), context: child.userContext, payload, error: serializeError(error) });
|
||||
throw error;
|
||||
}
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
## Runtime adapters
|
||||
|
||||
Core packages should not import Node-only APIs.
|
||||
|
||||
Possible implementations:
|
||||
|
||||
- Node adapter: `AsyncLocalStorage` for context, optional `diagnostics_channel` publishing.
|
||||
- Browser/workers fallback: local subscriber set and limited/manual context propagation.
|
||||
- Bun/Deno adapters: use runtime-specific async context if available.
|
||||
|
||||
For Node, diagnostics channels can be used as a passive event bus:
|
||||
|
||||
```ts
|
||||
import { channel } from "diagnostics_channel";
|
||||
channel("pi.observability").publish(event);
|
||||
```
|
||||
|
||||
Subscribers can create OTel/Sentry spans without monkey-patching pi.
|
||||
|
||||
## What pi emits
|
||||
|
||||
Pi emits what happened. It does not create OTel/Sentry spans directly.
|
||||
|
||||
Initial minimal event names:
|
||||
|
||||
```text
|
||||
pi.agent.prompt
|
||||
pi.agent.skill
|
||||
pi.agent.prompt_template
|
||||
pi.agent.compaction
|
||||
pi.agent.branch_navigation
|
||||
pi.agent.session.append_entry
|
||||
pi.ai.provider.request
|
||||
```
|
||||
|
||||
Each operation emits:
|
||||
|
||||
```text
|
||||
start
|
||||
end
|
||||
error
|
||||
```
|
||||
|
||||
Later additions:
|
||||
|
||||
```text
|
||||
pi.agent.turn
|
||||
pi.agent.tool_call
|
||||
pi.agent.queue_update
|
||||
pi.ai.provider.retry
|
||||
pi.ai.provider.first_token
|
||||
pi.ai.provider.usage
|
||||
pi.session.read
|
||||
pi.session.write
|
||||
```
|
||||
|
||||
## Minimal instrumentation points
|
||||
|
||||
### packages/agent
|
||||
|
||||
Wrap:
|
||||
|
||||
- `AgentHarness.prompt()`
|
||||
- `AgentHarness.skill()`
|
||||
- `AgentHarness.promptFromTemplate()`
|
||||
- `AgentHarness.compact()`
|
||||
- `AgentHarness.navigateTree()`
|
||||
- `Session.appendTypedEntry()` or storage append facade
|
||||
|
||||
Example:
|
||||
|
||||
```ts
|
||||
return traceOperation(
|
||||
"pi.agent.prompt",
|
||||
{
|
||||
sessionId: turnState.sessionId,
|
||||
provider: turnState.model.provider,
|
||||
model: turnState.model.id,
|
||||
promptLength: text.length,
|
||||
imageCount: options?.images?.length ?? 0,
|
||||
},
|
||||
() => this.executeTurn(turnState, text, options),
|
||||
);
|
||||
```
|
||||
|
||||
Session write:
|
||||
|
||||
```ts
|
||||
return traceOperation(
|
||||
"pi.agent.session.append_entry",
|
||||
{ entryType: entry.type },
|
||||
async () => {
|
||||
await this.unwrap(this.storage.appendEntry(entry));
|
||||
return entry.id;
|
||||
},
|
||||
);
|
||||
```
|
||||
|
||||
### packages/ai
|
||||
|
||||
Wrap common provider boundaries:
|
||||
|
||||
- `streamSimple()`
|
||||
- `completeSimple()`
|
||||
|
||||
Example:
|
||||
|
||||
```ts
|
||||
return traceOperation(
|
||||
"pi.ai.provider.request",
|
||||
{
|
||||
api: model.api,
|
||||
provider: model.provider,
|
||||
model: model.id,
|
||||
sessionId: options.sessionId,
|
||||
reasoning: options.reasoning,
|
||||
},
|
||||
() => actualStreamSimple(model, context, options),
|
||||
);
|
||||
```
|
||||
|
||||
End/error payloads can include safe metadata:
|
||||
|
||||
- stop reason
|
||||
- status code
|
||||
- retry count
|
||||
- input/output/total tokens
|
||||
- cost total
|
||||
- aborted/timeout flag
|
||||
|
||||
## Safety and redaction
|
||||
|
||||
Default payloads must be safe.
|
||||
|
||||
Safe by default:
|
||||
|
||||
- provider
|
||||
- model
|
||||
- API identifier
|
||||
- session id
|
||||
- entry type
|
||||
- tool name
|
||||
- status code
|
||||
- stop reason
|
||||
- token counts
|
||||
- costs
|
||||
- durations
|
||||
|
||||
Unsafe by default:
|
||||
|
||||
- prompts
|
||||
- completions
|
||||
- tool args
|
||||
- tool results
|
||||
- shell output
|
||||
- file contents
|
||||
- provider request payloads
|
||||
- provider response bodies
|
||||
- API keys
|
||||
- headers
|
||||
|
||||
Content capture can be opt-in later with explicit redaction hooks.
|
||||
|
||||
## Listener behavior
|
||||
|
||||
Observability must never affect pi execution.
|
||||
|
||||
Subscriber errors should be swallowed or isolated. Harness hooks are control-plane and may affect execution; observability subscribers are passive and must not.
|
||||
|
||||
## User context
|
||||
|
||||
Users can associate arbitrary context with a turn:
|
||||
|
||||
```ts
|
||||
await runWithPiContext(
|
||||
{
|
||||
userId: "u123",
|
||||
orgId: "acme",
|
||||
region: "eu",
|
||||
},
|
||||
() => harness.prompt("fix this"),
|
||||
);
|
||||
```
|
||||
|
||||
Every emitted event inside that async chain includes the context:
|
||||
|
||||
```ts
|
||||
{
|
||||
type: "start",
|
||||
name: "pi.ai.provider.request",
|
||||
traceId: "t1",
|
||||
spanId: "s3",
|
||||
parentSpanId: "s1",
|
||||
context: {
|
||||
userId: "u123",
|
||||
orgId: "acme",
|
||||
region: "eu",
|
||||
},
|
||||
payload: {
|
||||
provider: "anthropic",
|
||||
model: "claude-sonnet-4",
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
An OTel adapter can map this to span attributes. A Sentry adapter can map it to Sentry context/spans. A custom user can log JSON.
|
||||
|
||||
## Package story
|
||||
|
||||
Minimal initial package:
|
||||
|
||||
```text
|
||||
packages/observability
|
||||
runtime-agnostic context + traceOperation + subscribe
|
||||
```
|
||||
|
||||
Then:
|
||||
|
||||
```text
|
||||
packages/ai
|
||||
emits pi.ai.* events
|
||||
|
||||
packages/agent
|
||||
emits pi.agent.* / pi.session.* events
|
||||
```
|
||||
|
||||
Optional later:
|
||||
|
||||
```text
|
||||
packages/observability-node
|
||||
AsyncLocalStorage + diagnostics_channel bridge
|
||||
|
||||
packages/otel
|
||||
subscribes to pi events and creates OpenTelemetry spans
|
||||
```
|
||||
|
||||
## Thesis
|
||||
|
||||
Pi defines a stable, safe event contract. Adapters define where events go.
|
||||
|
||||
This makes ai/harness observable without binding core packages to OTel, Sentry, Node-only APIs, or monkey-patching.
|
||||
Reference in New Issue
Block a user