OMP Workbook

Read the source. Follow the evidence.

Debugging and verification

A workbook milestone is complete when the intended behavior is observable at the right layer.

A successful import is not proof of a useful tool. A successful tool adapter call is not proof that the model chooses it. A headless no-op is not proof of a TUI.

A practical debugging order

“The extension did not load”

Check:

  1. The launch selected the intended file or directory.
  2. The module has a default factory.
  3. Adjacent assets such as tool.txt exist.
  4. Runtime dependencies match the custom host.
  5. The file is not hidden/ignored by the discovery route.
  6. A disabled derived ID or plugin state did not filter it.
  7. An earlier same-name capability item did not shadow it.

Inspect the path-specific loader error. Do not replace a failing factory with an empty function merely to obtain “loaded.”

Terminal shell—default-profile log example; use the actual active state-root log path on your host:

Debugging and verification · source excerpt 1; read surrounding instructions
tail -f ~/.omp/logs/omp.$(date +%F).*.log

The log root can differ with host/state-root configuration. The supplied startup watchdog prints the actual log path.

Terminal shell—source-backed startup diagnostics:

Debugging and verification · source excerpt 2; read surrounding instructions
PI_DEBUG_STARTUP=1 omp --no-extensions -e "$EXAMPLES/package-lab/single/field-notes.ts"

This reveals startup phases; it is not an isolation flag.

“The command works, but the agent cannot use it”

Inspect the tool registry.

  • Was a tool actually registered?
  • Is it enabled?
  • Is it discoverable rather than top-level?
  • Does its description explain operations and prerequisites?
  • Are important result fields in model-visible content?
  • Is the agent trying to invoke a slash command?

Use a direct schema-valid tool scenario first. Then, separately, test a real model session if useful.

“Enter keeps completing instead of submitting”

Test exact completed input, not only partial input.

  • Completion values should contain the full argument text.
  • A sole exact completed match should return null.
  • Test Enter with the completion menu in the state the user actually encounters.
  • Do not always press Escape first and then claim the exact-match behavior is proven.

“The UI is blank”

Check mode, not just hasUI.

  • RPC has semantic UI but no native component factory.
  • ACP needs negotiated form support.
  • Print/JSON default methods are inert.
  • setFooter and setHeader are no-op in the supplied TUI implementation.
  • Theme objects are not accepted by the supplied TUI setter.
  • A tool without a renderer should retain its normal content path.

For a custom component, inspect width bounds, focus, abort, disposal and restoration of the editor.

“State leaked into another session”

Ask which lifetime owns it.

  • Module-scope mutable variable?
  • Factory-local closure reused after /new?
  • Whole-journal scan instead of getBranch()?
  • Stale context cached across a workspace move?
  • Parent-bound extension instances passed to a new SDK host?
  • Background sendMessage() sent to the current runtime instead of a captured owner?

The Field Notes result is not a leak according to its contract: its selection is binding-local. Calling it session-local would be the bug in the explanation.

“Permission came back after revocation”

Look for an in-flight dialog.

Clearing the current grant is insufficient. Abort the presentation and invalidate its generation. A late positive answer must fail the generation check.

Test the delayed positive response deliberately; do not test only revoke-after-grant.

“Reload ignored my code edit”

Determine which reload path ran.

A session reload is not a factory rebind. Restart the explicit launch, then inspect a changed description or deterministic behavior. Do not use an old closure’s output as evidence that new source was imported.

“A result arrived twice—or in the wrong branch”

Inspect:

  • captured target;
  • namespace and delivery ID;
  • owner/session file lineage;
  • anchor presence on the active branch;
  • reset boundaries;
  • receipt state versus wake state.

Retry a deferred/pending delivery with the same identity and body. Do not recapture the active session merely because the original owner is inactive.

“A provider is listed but unusable”

Separate:

  • catalog registration;
  • cached discovery;
  • available/auth-configured selection;
  • credential resolution;
  • streaming transport;
  • final inference.

A catalog or rollback test does not prove OAuth or provider behavior.

“A denied write never reaches the fallback”

Check whether it is actually:

  • a supported ordinary byte-write/unlink;
  • a permission error;
  • a canonicalizable destination;
  • inside an initialized fallback lifecycle;
  • allowed by the handler’s session/path policy.

Archives, SQLite, subprocesses and remote ACP writes do not enter this seam.

Verification levels

LevelWhat it provesWhat it does not prove
Schema/type checkingPublic API compatibility and parameter/result typingRuntime discovery or useful behavior
Pure domain testsTransition rules, validation and data invariantsHost wiring
Real loader/runner/adapter scenarioActual binding, interception and event contractsModel choice or physical UI
Real session storage scenarioJournal/reopen/branch behaviorPower-loss durability or cross-process coordination
Protocol roundtripRequests, responses, cancellation and degraded modesA particular remote client’s visual/accessibility quality
Real TUI/controller/composer smokeActual focus, keys, rendering and editor preservationEvery physical terminal or IME
Real provider/service integrationActual endpoint/auth/transport behaviorGeneral safety for unrelated providers/configurations

A small test you can run without a provider

The Review Desk domain is a useful offline test target. This is a new exercise test file, not one of the downloaded proof runners.

Complete TypeScript test exercise—save beside Review Desk’s domain.ts as domain.test.ts:

Debugging and verification · source excerpt 3; read surrounding instructions
import { expect, test } from "bun:test";
import { MAX_TEXT_LENGTH, transition } from "./domain";

test("accept keeps edited text and increments once", () => {
    const before = { revision: 0, status: "draft" as const, text: "Original." };
    expect(transition(before, "accept", "Reviewed.")).toEqual({
        revision: 1,
        status: "accepted",
        text: "Reviewed.",
    });
});

test("reject and cancel preserve pre-review text", () => {
    const before = { revision: 4, status: "draft" as const, text: "Keep this." };
    expect(transition(before, "reject").text).toBe("Keep this.");
    expect(transition(before, "cancel")).toEqual({
        revision: 5,
        status: "cancelled",
        text: "Keep this.",
    });
});

test("invalid replacement text fails", () => {
    const before = { revision: 0, status: "draft" as const, text: "Original." };
    expect(() => transition(before, "revise", "   ")).toThrow();
    expect(() => transition(before, "revise", "x".repeat(MAX_TEXT_LENGTH + 1))).toThrow();
    expect(before.revision).toBe(0);
});

Terminal shell—from the downloaded Review Desk directory:

Debugging and verification · source excerpt 4; read surrounding instructions
bun test ./domain.test.ts

Expected checkpoint: the three domain tests pass without model inference. They do not test the permission-dialog race; that requires the real runtime/protocol scenario.

Story-level acceptance scenarios

Use these as falsifiable checks:

Seed Desk

  • Welcome command and tool return the same hours.
  • Exact completion does not trap Enter.
  • Herb query returns only basil with eight fixture packets.
  • Unknown ID fails.
  • Cancelled human grant/reservation appends no state.
  • Agent act before grant fails.
  • Successful act advances revision.
  • Old revision fails without a second mutation.
  • Reopen restores state.
  • Sibling branches do not mix reservations.
  • New session starts with empty reservations and authority off.

Review Desk

  • Fresh inspection reports revision 0 and no grant.
  • Accept stores the edited note once.
  • Reject preserves pre-dialog text.
  • Editor cancellation records a distinct cancelled outcome.
  • Stale agent act fails.
  • Successful act consumes the grant.
  • Revocation during a pending dialog emits cancellation and defeats a late positive response.
  • Overlay cleanup preserves composer text.
  • Print mode refuses interactive changes.
  • RPC dialogs work without invoking native component factories.
  • ACP with and without form capability has explicitly tested behavior.

Package Lab

  • Single file, index directory and manifest resolve to the expected entries.
  • Helpers are not accidentally bound as extensions.
  • Optional feature defaults off.
  • Explicit optional entry works separately.
  • Invalid selection preserves the earlier selection.
  • /new and switching preserve binding-local selection.
  • A new binding resets it.
  • Resource dispatch preserves provenance.
  • Factory failure does not suppress later entries.
  • Provider rollback is not mistaken for rollback of all side effects.

The optional local extb toolkit

The supplied local extension-builder toolkit is separate from native OMP. It is not included among this workbook’s public example download paths, and readers should not assume a particular developer’s local installation exists.

Its new, mock, smoke and install commands illustrate a useful process, but its implementation has a narrower scope:

  • The mock API implements command registration, not the full ExtensionAPI.
  • Its factory call does not await async initialization.
  • It cannot validate these async, tool-registering examples as a general OMP loader.
  • Its smoke driver uses a real PTY but inherits environment/configuration and selects a default model.
  • It presses Escape before Enter, so it does not independently prove the exact-match-menu case.
  • Its 25-second subprocess timeout is a template choice, not proof that native command handlers have a universal 30-second budget.
  • Its global install path is a fixed copy target, not a complete profile-aware installer.
  • Plain extensions/ needs explicit loading; native .omp/extensions/ does not.

Use the toolkit’s process idea—small scaffold, domain test, real host smoke, deliberate installation—without treating its hints as stronger than the current runtime.

Source, snapshot 2026-08-29: supplied contract tests and observed report; extension-builder/bin/extb.ts, harness/mock-pi.ts, harness/tui-smoke.ts; packages/coding-agent/src/main.ts; linked public examples.

Extensions inside those boundaries · Source chapter: extensions/debugging-and-verification. Original evidence remains scoped to its recorded snapshot.

Read this chapter as Markdown

Your lesson ticks

A self-reported reading checklist, not proof of real OMP behavior. Only these ticks are saved in this browser. Reading a milestone does not resume, fork, reset or export a session.

Chapters I have worked through
Start here 1
Sessions, resets, and reviewable history 19
Memory and reusable knowledge 14
Tangent work and live control 17
Tool permissions and approvals 15
Extensions inside those boundaries 23
Connections and next steps 8
0 of 97 checked

Checklist saving needs JavaScript and available browser storage.