## Context `docs/ROADMAP.md` places the Planner inside Milestone 3 (Agent Runtime): "Planner/Executor structure, tool execution, retry behavior, and task runner," already implemented by `apex-agent-mvp`. What `apex-agent-mvp` actually shipped for the Planner half is `runtime/planner.py::Planner` — a stub that returns one hardcoded `describe_screen` step, then `[]` forever. Every later milestone built real capability around this stub (`Executor` retry/backoff, `TaskContext` memory, `semantic-scene` enrichment, `world-model` cross-step state) without ever giving the loop a real decision -maker. This change closes that gap: an `AIPlanner` that uses native LLM tool/function calling to choose one grounded action per turn, wired into `TaskRunner` behind a default-off flag so applying this change does not change any existing task's behavior, cost, or dependencies unless explicitly enabled. Two decisions were fixed before design work started (not re-litigated here): a **pluggable dual-provider** abstraction (Anthropic tool use and OpenAI function calling, switchable via config) rather than a single hardcoded provider, and **Scene + screenshot** (vision multimodal) as the Planner's perception input rather than Scene-only — the latter requires the Constitution amendment covered in Decision D5. ## Goals / Non-Goals **Goals:** - Replace the stub decision logic with a real LLM call that chooses exactly one tool (action or `finish_task`) per `plan()` invocation, using the current `Scene`, optionally the current screenshot, and recent `WorldState` history as input. - Support both Anthropic and OpenAI as interchangeable providers behind one internal `ToolCallingClient` interface, selected by configuration alone. - Keep the change strictly additive and default-disabled: with `AI_PLANNER_ENABLED` unset (or false), `TaskRunner`'s constructed planner, control flow, and dependencies are byte-for-byte the same as before this change. - Make task completion and task failure both explicit, model-driven signals (`finish_task(success, reason)`) rather than inferring completion from the absence of a tool call. - Preserve every Constitution invariant except the one explicitly amended (Perception Boundary), and amend that one as narrowly as the vision requirement allows. **Non-Goals:** - No multi-step lookahead planning — `AIPlanner.plan()` always returns 0 or 1 step; `TaskRunner.run()` already re-observes and re-plans every loop iteration, so pre-planning multiple steps ahead would only drift from the actual screen state. - No `wait`/no-op tool. Forcing "exactly one tool call per turn" means the model cannot express "do nothing this turn" — a real but low-severity gap (a spurious action self-corrects next turn once the loop re-observes). Candidate for a v1.1 follow-up, not this change. - No dynamic tool subset per task or app — the six-tool face (`ACTION_TOOL_SPECS` + `finish_task`) is fixed. - No `find_text_on_screen`/`find_icon_on_screen`/`describe_screen_semantic` exposed as Planner tools — the per-turn `Scene` JSON already enumerates every element's id/type/text/bounds, so a "search for X" tool call would only add latency and hallucination surface without adding information the model doesn't already have. - No automatic retry of a failed LLM call inside `AIPlanner` or `ToolCallingClient` — a transient failure surfaces as a failed task via the `TaskRunner.run()` fix in D7, matching the project's existing convention that retry is `Executor`'s concern for tool execution, not the Planner's concern for its own decision calls. - No provider/model selection exposed as an API request parameter — v1 is a deployment-time/environment-variable choice only. - No backfill of `openspec/specs/agent-runtime/`, which does not exist because `apex-agent-mvp`'s `agent-runtime` capability was never archived into `openspec/specs/`. This is a pre-existing gap unrelated to this change; this change's `specs/agent-runtime/spec.md` uses `## ADDED Requirements` against that unarchived baseline, the same way `semantic-scene-runtime` and `world-model-runtime` each did. ## Decisions ### D1: Single-step, ReAct-style `plan()` — never multi-step `AIPlanner.plan()` returns a list of 0 or 1 `PlannedStep`. `TaskRunner.run()` already calls `observer()` then `_plan()` fresh on every loop iteration (`runtime/task.py`), so a Planner that tried to hand back several steps at once would either have those extra steps silently ignored by the loop's current per-iteration contract, or require changing that contract to consume a queue — both worse than just re-deciding every turn against the freshly observed `Scene`. **Alternative considered**: return a short queued plan (e.g. up to 3 steps) and only re-plan when a step's expectation is violated. Rejected — UI state can change after any single action (a dialog appears, a keyboard covers an element), so a queued step is frequently stale by the time it would execute; single-step keeps every action grounded in the turn's actual `Scene`. ### D2: Explicit `finish_task(success, reason)` control tool, not an implicit "no call" signal Task completion and task failure are both signaled by the model calling `finish_task` — never by the model declining to call any tool (which native tool-calling APIs do not reliably support as a distinguishable "done" state across both providers) and never by a heuristic on the Planner's side. `finish_task(success=True)` maps to `AIPlanner.plan()` returning `[]`, which reuses `TaskRunner.run()`'s existing `if not steps or ...: return self._complete_task(task)` short-circuit unchanged. `finish_task(success= False, reason)` maps to `AIPlanner.plan()` raising `TaskFailedError(reason)` — reusing `core/errors.py`'s `TaskFailedError`, defined since `apex-agent-mvp` but never previously constructed anywhere in the codebase. These two outcomes are deliberately routed through different mechanisms (empty list vs. exception) rather than both returning `[]` with a status flag, because `run()`'s short-circuit condition is `if not steps or self.planner.goal_reached(...)` — a bare empty list on failure would be silently read as success by that exact line. **Alternative considered**: signal success via `goal_reached()` returning `True`. Rejected — see D3. ### D3: `AIPlanner.goal_reached()` always returns `False` This is a required invariant, not a style choice. `run()` evaluates `if not steps or self.planner.goal_reached(...)` — the `or` means `goal_reached()` is still consulted even when `steps` is non-empty. If `goal_reached()` could ever return `True` on a turn where `plan()` also returned a real action, that action would be silently discarded and the task would be marked complete one turn early. Routing every completion/failure signal exclusively through `finish_task` (D2) removes any reason for `goal_reached()` to do anything; it is a permanent no-op override, documented as such at the call site. ### D4: Dual-provider abstraction via `Protocol`, forced single-tool-call on both `runtime/tool_calling_client.py` defines `ToolCallingClient` as a structural `Protocol` (matching the existing precedent of `semantic/enricher.py::SemanticLLMClient` and `skills_learning/embeddings.py::EmbeddingClient` — no ABC is used anywhere in this codebase for a swappable single-method client), with concrete `AnthropicToolCallingClient` and `OpenAIToolCallingClient` implementations selected by `build_client(config)` based on `AI_PLANNER_PROVIDER`. Both implementations force the API to return exactly one tool call per turn — Anthropic via `tool_choice={"type": "any", "disable_parallel_tool_use": True}`, OpenAI via `tool_choice="required"` plus the top-level `parallel_tool_calls=False` — so a single `ToolCallDecision` can always be parsed deterministically regardless of provider. Both implementations use the lazy-import-plus-injectable-transport pattern already established by `semantic/llm_client.py::AnthropicSemanticClient` (SDK imported only inside a `_client()` method; constructor accepts an optional `transport` for tests), and wrap every SDK/network/parse failure into one internal `ToolCallUnavailable` exception, mirroring `EnrichmentUnavailable`. **Alternative considered**: a single hardcoded Anthropic-only client (the simpler default for a first LLM-driven Planner). Explicitly rejected by product decision before this design was drafted — dual-provider pluggability is a requirement, not a nice-to-have, for this change. ### D5: Perception Boundary amendment — narrow, Planner-only screenshot exception The Constitution's Perception Boundary previously stated `Scene` is the *only* perception artifact the LLM sees. Vision-grounded action selection (precise tap/swipe coordinates, disambiguating visually-similar elements that OCR/UI-tree fusion can conflate) requires the Planner to also see the raw screenshot for the current step — a deliberate product decision, not an oversight. The amendment is written to be as narrow as that requirement actually is: it names one consumer (`the runtime-layer AI Planner, and only that Planner`) and explicitly reaffirms that every other layer/consumer (`api`, `tools`, `perception`, `storage`, or any other LLM consumer) still never receives raw screenshot bytes, and that `Scene` itself is still produced exclusively through `PerceptionProvider`. See "Constitution Compliance" below for the exact before/after text. **Alternative considered**: keep `Scene` as the Planner's only input (Scene-only, no vision). Explicitly rejected by product decision before this design was drafted, for the reasons above. **Alternative considered**: broaden the exception to "any runtime-layer LLM consumer" instead of naming the Planner specifically, anticipating future vision consumers. Rejected — Change Discipline requires each change to justify its own scope; a hypothetical future consumer should justify its own amendment when it exists, not inherit a pre-approved blank check today. ### D6: Default-disabled via `AI_PLANNER_ENABLED`, mirroring existing LLM-feature flags `PlannerConfig.enabled` defaults to `False`, and `load_config()` reads `AI_PLANNER_ENABLED` the same way `semantic/config.py`, `skills_learning/config.py`, and `agents/config.py` gate their own LLM-backed behavior — off unless explicitly turned on. (`world/config.py`'s `WorldConfig.enabled` defaults `True`, but that capability makes zero LLM calls and is not a counterexample to this convention.) `TaskRunner`'s `_default_planner()` follows the same "self-loaded config, enabled flag picks the implementation" shape already used for `world_model` and `on_task_succeeded` (`runtime/task.py`), so no change to `api/rest.py` or any other caller of `TaskRunner()` is required to preserve current behavior. ### D7: `TaskRunner.run()` gains a try/except around observe+plan Before this change, `run()` had no error handling around `self.observer(...)` or `self._plan(...)` — safe only because the stub `Planner` never raised. `AIPlanner` can raise for real reasons (network failure, malformed provider response, `TaskFailedError` from D2's `finish_task(success=False)` path), so `run()` now catches any exception from that per-iteration observe-then-plan sequence and marks the task `status="failed"` with `failure_reason=f"{type(exc).__name__}: {exc}"`, returning immediately. Two failure modes are avoided by doing this explicitly rather than skipping it: an uncaught exception leaving a `BackgroundTasks`-invoked task stuck at `status="running"` forever, and a catch-and-return-`[]` degrade pattern (as used by `semantic/enricher.py::enrich_scene()`) that would be misread as task success by the same `if not steps or ...` short-circuit discussed in D2/D3. **Alternative considered**: copy `enrich_scene()`'s catch-and-degrade-to-None pattern. Rejected — that pattern is correct for an optional enrichment pass where the caller has a defined fallback (use the raw `Scene`). The Planner has no such fallback: if it cannot decide, the task cannot proceed, and pretending otherwise (via an empty list) means silent, incorrect "success." ### D8: `_planner_accepts_world()` generalized to `_planner_accepts(name)` `TaskRunner` already used `inspect.signature` reflection to decide whether to pass a `world=` kwarg to `Planner.plan()`, so that narrow-signature test Planners (e.g. `NoWorldPlanner`) would not receive an unexpected keyword argument. This change generalizes that one method to take the parameter name as an argument and reuses it for both `"world"` and the new `"screenshot"` kwarg, preserving backward compatibility with any existing `Planner` subclass that does not declare a `screenshot` parameter (or `**kwargs`). ## Constitution Compliance - **Device Boundary**: unaffected — this change adds no driver code and does not touch `driver/`/`device/`. - **Tool Boundary**: unaffected — `AIPlanner` calls tools exclusively through `PlannedStep` → `Executor`, the same as the stub `Planner`; no new code imports a concrete driver or bypasses `tools/`. - **Perception Boundary**: amended, narrowly, per D5. The invariant that `Scene` is the only perception artifact for every consumer *except* the named Planner is preserved and made explicit in the same sentence as the exception, rather than being loosened globally. - **Runtime Boundary**: preserved and, in fact, fulfilled by this change — the existing text already states "LLM dependencies enter at `runtime` through Planner behavior," which is exactly what `AIPlanner` is. The Executor remains the only component that calls tools and handles tool-execution retries; `AIPlanner` does not call tools directly or implement its own retry loop (D7 fails the task instead of retrying). - **Change Discipline**: the dependency direction (`core -> driver/device -> tools -> perception -> storage -> runtime -> api`) is preserved — all new code lives in `runtime/`, imports flow downward only (`runtime/ai_planner.py` imports from `runtime/`, `core/`; nothing in `core`, `driver`, `device`, or `tools` imports anything LLM-related). The new external integration (the OpenAI SDK, newly *used* though already declared as a dependency) is added at the adapter layer that owns the concern (`runtime/tool_calling_client.py`), not at the domain or device boundary. ## Risks / Trade-offs - **[Risk] Per-step LLM call adds latency and cost to every planning turn** → Mitigation: default-disabled (D6); `AI_PLANNER_TIMEOUT_SECONDS` bounds worst-case latency; no in-`AIPlanner` retry (D7) means a slow/failing provider fails fast instead of compounding delay. - **[Risk] Hallucinated tool arguments (e.g. tap coordinates outside any element's bounds)** → Mitigation: the system prompt instructs the model to ground every coordinate in the current turn's `Scene` element bounds (and screenshot, when present); a bad tap still degrades gracefully into an ordinary failed/retried step through `Executor`'s existing mechanism, unchanged by this design. - **[Risk] Provider wire-format drift (Anthropic/OpenAI SDK or API changes) silently breaks request construction or response parsing** → Mitigation: all format-specific logic is isolated into small, independently unit- tested functions (`_anthropic_tool`, `_openai_tool`, `_decision_from_anthropic_response`, `_decision_from_openai_response`) against fake transports, so a drift shows up as a specific, localized test failure rather than a silent behavior change. - **[Risk] Two providers could behave inconsistently** (e.g. one honors `finish_task` semantics more reliably than the other) → Mitigation: both are constrained by the identical `ALL_TOOL_SPECS` JSON Schema and the same system prompt; provider-specific behavior differences are a model-quality concern to observe via the optional real-integration test, not something this design can fully eliminate structurally. - **[Risk] Perception Boundary amendment could be read as a precedent for loosening the invariant further** → Mitigation: the amendment text itself names exactly one consumer and restates the "no one else" constraint in the same breath (D5); any future consumer needs its own amendment. ## Migration Plan Purely additive; no data migration: 1. Add the five new `runtime/*` files and the `runtime/planner.py`/ `runtime/task.py` edits described in Impact. With `AI_PLANNER_ENABLED` unset, `TaskRunner()` continues to construct the stub `Planner` exactly as before. 2. Amend `docs/CONSTITUTION.md`'s Perception Boundary section (D5). 3. No changes to `storage/`'s timeline format, `api/rest.py`, or `api/mcp.py`. 4. Rollback is simply leaving `AI_PLANNER_ENABLED` unset/false, or reverting the changed files; no other capability depends on this one. 5. Enabling in a real environment requires setting `AI_PLANNER_ENABLED=true`, `AI_PLANNER_PROVIDER` (`anthropic` or `openai`), and the corresponding provider's API key in the process environment (already-existing SDK convention, not a new config surface this change introduces). ## Open Questions - Whether a `wait`/no-op tool is actually needed in practice, or whether the self-correcting-next-turn behavior of a spurious action is good enough indefinitely — deferred to real usage observation, not decided here. - Whether `AI_PLANNER_TIMEOUT_SECONDS`'s default (30s) is well-tuned against real provider latency for image-bearing requests — needs tuning against real usage data once this is enabled somewhere with real traffic; not fixed by this design. - Whether provider/model selection should eventually move from environment-variable/deployment-time to a per-task or per-request choice — leaning toward "not until a concrete need appears" (YAGNI), consistent with `semantic-scene-runtime`'s D4 reasoning for its own model choice.