Boundary Desk: one call at a time
Boundary Desk begins with a deceptively small question: what did the word Approve authorize?
For the generic tool gate in the supplied wrapper, the answer is the current call. There is no built-in choice here named “allow for session” or “always allow,” and selecting Approve does not write a per-tool policy.
Inspect the actual binary choice
The wrapper calls:
choice = await uiContext.select(safetyPrompt, ["Approve", "Deny"]);
const approved = choice === "Approve";
Only the exact returned label Approve is positive. Deny, undefined, and other nonmatching values are not approval. A dialog exception also stops this path and is reported as a false resolution before being rethrown.
These are the generic tool-approval choices. They are not Review Desk’s local review outcomes, an ACP plan-review menu, or a domain delegation dialog.
Predict the second call
Inspect boundary-approval-does-not-persist:
{
"approval": "exec",
"args": { "action": "publish" },
"mode": "always-ask",
"choices": ["Approve", "Deny"],
"repeat": 2
}
Prediction: does the first answer eliminate the second question?
Recorded answer: no. The same wrapped tool, with unchanged settings, was invoked twice. There were two prompts and two tool_call events, but only one successful call and one inert executor entry. The second invocation threw Tool call denied by user: seed_note.
The first inert result remains present in the report even though the final outcome is a throw. That is not evidence that the second invocation ran. Read successfulCalls, the execution ledger, and the per-call lifecycle together.
The simpler cases establish the alternatives independently:
| Case | Recorded response | Prompts | Inert executions | Outcome |
|---|---|---|---|---|
boundary-approve-one-call | Approve | 1 | 1 | return |
boundary-deny-one-call | Deny | 1 | 0 | throw |
boundary-dismiss-one-call | undefined, represented by fixture null | 1 | 0 | throw |
A declined call is not a persisted deny policy. Conversely, a user-policy deny does not require a dialog in which a person declines this particular call.
Follow the order, not just the final banner
For the directly exercised wrapper path, the useful sequence is:
- Consume any existing loop-emission marker and sample approval inputs from the execute-time context.
- Resolve the original input; an effective deny stops this wrapper path before its own
tool_callhandler emission. - If not already emitted by the loop, run applicable
tool_callhandlers, which can block or return replacement input. - Resolve approval against the effective input and report an effective xdev tier when that callback exists.
- If approval is required, wait for the applicable scheduled-call preview when such a waiter is installed.
- Emit
tool_approval_requestedwhen approval lifecycle handlers are present. - Obtain the selection, or resolve false because no UI is available.
- Emit
tool_approval_resolved; execute the underlying tool only after a positive required decision.
ui.select and the executor are steps in this explanation, not extra OMP event names. The existing extensions-runner.test.ts records requested → UI selection → resolved order and tests preview waiting for canonical and wire-aliased tool names. The private fixture report records the lifecycle pairs, prompt data, and executor counts through its structural adapter; it does not exercise that transcript preview waiter.
An approval requested event can therefore exist even when no dialog was displayed. It records a required decision path, not proof of a human interaction. The empty sessionId in several private scenario events comes from their deliberately partial execute-time context, which has no session manager. It is not a real reader’s conversation identity.
Reclassification has a supported scope
The wrapper’s supported replacement path applies a returned input before the full approval gate. The existing runner tests verify that a revised input can become denied, that a prompt describes revised rather than original input, and that an xdev replacement forfeits the bypass.
For normal model-loop preparation, AgentLoopConfig.beforeToolCall documents replacement argument revalidation and propagation into scheduling and tool-call history. The retained Tools, interception and native delegation chapter explains the session wiring and its distinction from direct execution. A direct tool.execute() call is not proof that the loop’s schema revalidation occurred.
The runner’s handlers are not a field-by-field replacement pipeline: they receive the normalized event, and the last returned result object is retained unless a block short-circuits. Computer-provider calls expose a synthetic actions/safety-check view, so the wrapper does not apply ordinary returned input replacements to those execution parameters.
Do not extend the supported replacement proof to arbitrary in-process mutation, to code that changes data after approval, or to every nested executor. The wrapper samples mode and user policies before awaiting handlers and selection; it is not a continuous policy-revalidation transaction while a dialog remains open.
Cancellation is not one universal mechanism
The recorded dismissal case and synthetic RPC cancellation both produce undefined, which the wrapper refuses. Existing runner tests also cover a throwing approval selector and cancellation while an extension-owned tool_call confirmation is pending.
Those are different seams. The runner supplies signals and active-work timeout accounting to extension-handler dialogs. The generic approval selector call shown above does not itself pass dialog options containing the execution signal. Do not claim from these records that every external abort automatically cancels every native approval selection or defeats every possible late response. A cancelled presentation, a blocked tool, and rollback of an earlier effect remain separate claims.
Paper checkpoint: after the Approve-then-Deny repeat case, what is the execution count? Worked answer: one. The second prompt did not persist or reuse the first answer, and the denial did not undo the first inert record.
Failure boundary: the wrapper’s result hooks can transform returned content, details, and error presentation after execution. They cannot undo an external effect. An approval event, a nonthrowing result, or a visible success label is not a substitute for checking the operation’s own postcondition.
Source anchors: packages/coding-agent/src/extensibility/extensions/wrapper.ts — ExtensionToolWrapper.execute; packages/coding-agent/src/extensibility/extensions/runner.ts — emitToolCall, markToolCallEmitted, consumeToolCallEmitted, waitForToolApprovalPreview; packages/agent/src/types.ts — AgentLoopConfig.beforeToolCall; packages/coding-agent/test/extensions-runner.test.ts — tool approval lifecycle, input replacement, timeout, and cancellation tests.
Tool permissions and approvals · Source chapter: permissions/boundary-desk-one-call-at-a-time. Original evidence remains scoped to its recorded snapshot.