Files
agentic-mobile-control/openspec/changes/ai-planner-runtime/design.md
T
q792602257andClaude Sonnet 5 61ff3b425d feat(agent-runtime): add LLM-driven AI Planner with dual-provider tool calling
Replaces the stub Planner's fixed describe_screen/[] behavior with a real
decision-maker: AIPlanner uses native tool/function calling (Anthropic or
OpenAI, pluggable via AI_PLANNER_PROVIDER) to select exactly one grounded
action per turn, with an explicit finish_task(success, reason) tool for
completion/failure instead of an ambiguous "no tool call" signal. Default
disabled (AI_PLANNER_ENABLED=false) and additive; TaskRunner falls back to
the existing stub Planner unchanged when disabled.

Amends CONSTITUTION.md's Perception Boundary with one narrow exception:
only the AI Planner may receive the current step's raw screenshot bytes
alongside Scene, for vision-grounded coordinate grounding. Also fixes a
latent gap in TaskRunner.run(): observe/plan exceptions are now caught per
iteration and turned into a failed task with a failure_reason, instead of
propagating uncaught.

openspec change: ai-planner-runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-12 13:48:50 +08:00

290 lines
18 KiB
Markdown

## Context
`docs/ROADMAP.md` places the Planner inside Milestone 3 (Agent Runtime):
"Planner/Executor structure, tool execution, retry behavior, and task
runner," already implemented by `apex-agent-mvp`. What `apex-agent-mvp`
actually shipped for the Planner half is `runtime/planner.py::Planner` — a
stub that returns one hardcoded `describe_screen` step, then `[]` forever.
Every later milestone built real capability around this stub (`Executor`
retry/backoff, `TaskContext` memory, `semantic-scene` enrichment,
`world-model` cross-step state) without ever giving the loop a real decision
-maker. This change closes that gap: an `AIPlanner` that uses native LLM
tool/function calling to choose one grounded action per turn, wired into
`TaskRunner` behind a default-off flag so applying this change does not
change any existing task's behavior, cost, or dependencies unless explicitly
enabled.
Two decisions were fixed before design work started (not re-litigated here):
a **pluggable dual-provider** abstraction (Anthropic tool use and OpenAI
function calling, switchable via config) rather than a single hardcoded
provider, and **Scene + screenshot** (vision multimodal) as the Planner's
perception input rather than Scene-only — the latter requires the Constitution
amendment covered in Decision D5.
## Goals / Non-Goals
**Goals:**
- Replace the stub decision logic with a real LLM call that chooses exactly
one tool (action or `finish_task`) per `plan()` invocation, using the
current `Scene`, optionally the current screenshot, and recent
`WorldState` history as input.
- Support both Anthropic and OpenAI as interchangeable providers behind one
internal `ToolCallingClient` interface, selected by configuration alone.
- Keep the change strictly additive and default-disabled: with
`AI_PLANNER_ENABLED` unset (or false), `TaskRunner`'s constructed planner,
control flow, and dependencies are byte-for-byte the same as before this
change.
- Make task completion and task failure both explicit, model-driven signals
(`finish_task(success, reason)`) rather than inferring completion from the
absence of a tool call.
- Preserve every Constitution invariant except the one explicitly amended
(Perception Boundary), and amend that one as narrowly as the vision
requirement allows.
**Non-Goals:**
- No multi-step lookahead planning — `AIPlanner.plan()` always returns 0 or 1
step; `TaskRunner.run()` already re-observes and re-plans every loop
iteration, so pre-planning multiple steps ahead would only drift from the
actual screen state.
- No `wait`/no-op tool. Forcing "exactly one tool call per turn" means the
model cannot express "do nothing this turn" — a real but low-severity gap
(a spurious action self-corrects next turn once the loop re-observes).
Candidate for a v1.1 follow-up, not this change.
- No dynamic tool subset per task or app — the six-tool face
(`ACTION_TOOL_SPECS` + `finish_task`) is fixed.
- No `find_text_on_screen`/`find_icon_on_screen`/`describe_screen_semantic`
exposed as Planner tools — the per-turn `Scene` JSON already enumerates
every element's id/type/text/bounds, so a "search for X" tool call would
only add latency and hallucination surface without adding information the
model doesn't already have.
- No automatic retry of a failed LLM call inside `AIPlanner` or
`ToolCallingClient` — a transient failure surfaces as a failed task via the
`TaskRunner.run()` fix in D7, matching the project's existing convention
that retry is `Executor`'s concern for tool execution, not the Planner's
concern for its own decision calls.
- No provider/model selection exposed as an API request parameter — v1 is a
deployment-time/environment-variable choice only.
- No backfill of `openspec/specs/agent-runtime/`, which does not exist
because `apex-agent-mvp`'s `agent-runtime` capability was never archived
into `openspec/specs/`. This is a pre-existing gap unrelated to this
change; this change's `specs/agent-runtime/spec.md` uses `## ADDED
Requirements` against that unarchived baseline, the same way
`semantic-scene-runtime` and `world-model-runtime` each did.
## Decisions
### D1: Single-step, ReAct-style `plan()` — never multi-step
`AIPlanner.plan()` returns a list of 0 or 1 `PlannedStep`. `TaskRunner.run()`
already calls `observer()` then `_plan()` fresh on every loop iteration
(`runtime/task.py`), so a Planner that tried to hand back several steps at
once would either have those extra steps silently ignored by the loop's
current per-iteration contract, or require changing that contract to consume
a queue — both worse than just re-deciding every turn against the freshly
observed `Scene`.
**Alternative considered**: return a short queued plan (e.g. up to 3 steps)
and only re-plan when a step's expectation is violated. Rejected — UI state
can change after any single action (a dialog appears, a keyboard covers an
element), so a queued step is frequently stale by the time it would execute;
single-step keeps every action grounded in the turn's actual `Scene`.
### D2: Explicit `finish_task(success, reason)` control tool, not an implicit "no call" signal
Task completion and task failure are both signaled by the model calling
`finish_task` — never by the model declining to call any tool (which native
tool-calling APIs do not reliably support as a distinguishable "done" state
across both providers) and never by a heuristic on the Planner's side.
`finish_task(success=True)` maps to `AIPlanner.plan()` returning `[]`, which
reuses `TaskRunner.run()`'s existing `if not steps or ...: return
self._complete_task(task)` short-circuit unchanged. `finish_task(success=
False, reason)` maps to `AIPlanner.plan()` raising `TaskFailedError(reason)`
— reusing `core/errors.py`'s `TaskFailedError`, defined since
`apex-agent-mvp` but never previously constructed anywhere in the codebase.
These two outcomes are deliberately routed through different mechanisms
(empty list vs. exception) rather than both returning `[]` with a status
flag, because `run()`'s short-circuit condition is `if not steps or
self.planner.goal_reached(...)` — a bare empty list on failure would be
silently read as success by that exact line.
**Alternative considered**: signal success via `goal_reached()` returning
`True`. Rejected — see D3.
### D3: `AIPlanner.goal_reached()` always returns `False`
This is a required invariant, not a style choice. `run()` evaluates `if not
steps or self.planner.goal_reached(...)` — the `or` means `goal_reached()` is
still consulted even when `steps` is non-empty. If `goal_reached()` could
ever return `True` on a turn where `plan()` also returned a real action, that
action would be silently discarded and the task would be marked complete one
turn early. Routing every completion/failure signal exclusively through
`finish_task` (D2) removes any reason for `goal_reached()` to do anything;
it is a permanent no-op override, documented as such at the call site.
### D4: Dual-provider abstraction via `Protocol`, forced single-tool-call on both
`runtime/tool_calling_client.py` defines `ToolCallingClient` as a
structural `Protocol` (matching the existing precedent of
`semantic/enricher.py::SemanticLLMClient` and
`skills_learning/embeddings.py::EmbeddingClient` — no ABC is used anywhere in
this codebase for a swappable single-method client), with concrete
`AnthropicToolCallingClient` and `OpenAIToolCallingClient` implementations
selected by `build_client(config)` based on `AI_PLANNER_PROVIDER`. Both
implementations force the API to return exactly one tool call per turn —
Anthropic via `tool_choice={"type": "any", "disable_parallel_tool_use":
True}`, OpenAI via `tool_choice="required"` plus the top-level
`parallel_tool_calls=False` — so a single `ToolCallDecision` can always be
parsed deterministically regardless of provider. Both implementations use
the lazy-import-plus-injectable-transport pattern already established by
`semantic/llm_client.py::AnthropicSemanticClient` (SDK imported only inside a
`_client()` method; constructor accepts an optional `transport` for tests),
and wrap every SDK/network/parse failure into one internal
`ToolCallUnavailable` exception, mirroring `EnrichmentUnavailable`.
**Alternative considered**: a single hardcoded Anthropic-only client (the
simpler default for a first LLM-driven Planner). Explicitly rejected by
product decision before this design was drafted — dual-provider
pluggability is a requirement, not a nice-to-have, for this change.
### D5: Perception Boundary amendment — narrow, Planner-only screenshot exception
The Constitution's Perception Boundary previously stated `Scene` is the
*only* perception artifact the LLM sees. Vision-grounded action selection
(precise tap/swipe coordinates, disambiguating visually-similar elements
that OCR/UI-tree fusion can conflate) requires the Planner to also see the
raw screenshot for the current step — a deliberate product decision, not an
oversight. The amendment is written to be as narrow as that requirement
actually is: it names one consumer (`the runtime-layer AI Planner, and only
that Planner`) and explicitly reaffirms that every other layer/consumer
(`api`, `tools`, `perception`, `storage`, or any other LLM consumer) still
never receives raw screenshot bytes, and that `Scene` itself is still
produced exclusively through `PerceptionProvider`. See "Constitution
Compliance" below for the exact before/after text.
**Alternative considered**: keep `Scene` as the Planner's only input
(Scene-only, no vision). Explicitly rejected by product decision before this
design was drafted, for the reasons above.
**Alternative considered**: broaden the exception to "any runtime-layer LLM
consumer" instead of naming the Planner specifically, anticipating future
vision consumers. Rejected — Change Discipline requires each change to
justify its own scope; a hypothetical future consumer should justify its own
amendment when it exists, not inherit a pre-approved blank check today.
### D6: Default-disabled via `AI_PLANNER_ENABLED`, mirroring existing LLM-feature flags
`PlannerConfig.enabled` defaults to `False`, and `load_config()` reads
`AI_PLANNER_ENABLED` the same way `semantic/config.py`,
`skills_learning/config.py`, and `agents/config.py` gate their own
LLM-backed behavior — off unless explicitly turned on. (`world/config.py`'s
`WorldConfig.enabled` defaults `True`, but that capability makes zero LLM
calls and is not a counterexample to this convention.) `TaskRunner`'s
`_default_planner()` follows the same "self-loaded config, enabled flag
picks the implementation" shape already used for `world_model` and
`on_task_succeeded` (`runtime/task.py`), so no change to `api/rest.py` or any
other caller of `TaskRunner()` is required to preserve current behavior.
### D7: `TaskRunner.run()` gains a try/except around observe+plan
Before this change, `run()` had no error handling around
`self.observer(...)` or `self._plan(...)` — safe only because the stub
`Planner` never raised. `AIPlanner` can raise for real reasons (network
failure, malformed provider response, `TaskFailedError` from D2's
`finish_task(success=False)` path), so `run()` now catches any exception
from that per-iteration observe-then-plan sequence and marks the task
`status="failed"` with `failure_reason=f"{type(exc).__name__}: {exc}"`,
returning immediately. Two failure modes are avoided by doing this
explicitly rather than skipping it: an uncaught exception leaving a
`BackgroundTasks`-invoked task stuck at `status="running"` forever, and a
catch-and-return-`[]` degrade pattern (as used by
`semantic/enricher.py::enrich_scene()`) that would be misread as task
success by the same `if not steps or ...` short-circuit discussed in D2/D3.
**Alternative considered**: copy `enrich_scene()`'s catch-and-degrade-to-None
pattern. Rejected — that pattern is correct for an optional enrichment pass
where the caller has a defined fallback (use the raw `Scene`). The Planner
has no such fallback: if it cannot decide, the task cannot proceed, and
pretending otherwise (via an empty list) means silent, incorrect "success."
### D8: `_planner_accepts_world()` generalized to `_planner_accepts(name)`
`TaskRunner` already used `inspect.signature` reflection to decide whether to
pass a `world=` kwarg to `Planner.plan()`, so that narrow-signature test
Planners (e.g. `NoWorldPlanner`) would not receive an unexpected keyword
argument. This change generalizes that one method to take the parameter name
as an argument and reuses it for both `"world"` and the new `"screenshot"`
kwarg, preserving backward compatibility with any existing `Planner`
subclass that does not declare a `screenshot` parameter (or `**kwargs`).
## Constitution Compliance
- **Device Boundary**: unaffected — this change adds no driver code and does
not touch `driver/`/`device/`.
- **Tool Boundary**: unaffected — `AIPlanner` calls tools exclusively through
`PlannedStep``Executor`, the same as the stub `Planner`; no new code
imports a concrete driver or bypasses `tools/`.
- **Perception Boundary**: amended, narrowly, per D5. The invariant that
`Scene` is the only perception artifact for every consumer *except* the
named Planner is preserved and made explicit in the same sentence as the
exception, rather than being loosened globally.
- **Runtime Boundary**: preserved and, in fact, fulfilled by this change —
the existing text already states "LLM dependencies enter at `runtime`
through Planner behavior," which is exactly what `AIPlanner` is. The
Executor remains the only component that calls tools and handles
tool-execution retries; `AIPlanner` does not call tools directly or
implement its own retry loop (D7 fails the task instead of retrying).
- **Change Discipline**: the dependency direction (`core -> driver/device ->
tools -> perception -> storage -> runtime -> api`) is preserved — all new
code lives in `runtime/`, imports flow downward only (`runtime/ai_planner.py`
imports from `runtime/`, `core/`; nothing in `core`, `driver`, `device`, or
`tools` imports anything LLM-related). The new external integration (the
OpenAI SDK, newly *used* though already declared as a dependency) is added
at the adapter layer that owns the concern (`runtime/tool_calling_client.py`),
not at the domain or device boundary.
## Risks / Trade-offs
- **[Risk] Per-step LLM call adds latency and cost to every planning turn**
→ Mitigation: default-disabled (D6); `AI_PLANNER_TIMEOUT_SECONDS` bounds
worst-case latency; no in-`AIPlanner` retry (D7) means a slow/failing
provider fails fast instead of compounding delay.
- **[Risk] Hallucinated tool arguments (e.g. tap coordinates outside any
element's bounds)** → Mitigation: the system prompt instructs the model to
ground every coordinate in the current turn's `Scene` element bounds (and
screenshot, when present); a bad tap still degrades gracefully into an
ordinary failed/retried step through `Executor`'s existing mechanism,
unchanged by this design.
- **[Risk] Provider wire-format drift (Anthropic/OpenAI SDK or API changes)
silently breaks request construction or response parsing** → Mitigation:
all format-specific logic is isolated into small, independently unit-
tested functions (`_anthropic_tool`, `_openai_tool`,
`_decision_from_anthropic_response`, `_decision_from_openai_response`)
against fake transports, so a drift shows up as a specific, localized test
failure rather than a silent behavior change.
- **[Risk] Two providers could behave inconsistently** (e.g. one honors
`finish_task` semantics more reliably than the other) → Mitigation: both
are constrained by the identical `ALL_TOOL_SPECS` JSON Schema and the same
system prompt; provider-specific behavior differences are a model-quality
concern to observe via the optional real-integration test, not something
this design can fully eliminate structurally.
- **[Risk] Perception Boundary amendment could be read as a precedent for
loosening the invariant further** → Mitigation: the amendment text itself
names exactly one consumer and restates the "no one else" constraint in
the same breath (D5); any future consumer needs its own amendment.
## Migration Plan
Purely additive; no data migration:
1. Add the five new `runtime/*` files and the `runtime/planner.py`/
`runtime/task.py` edits described in Impact. With `AI_PLANNER_ENABLED`
unset, `TaskRunner()` continues to construct the stub `Planner` exactly as
before.
2. Amend `docs/CONSTITUTION.md`'s Perception Boundary section (D5).
3. No changes to `storage/`'s timeline format, `api/rest.py`, or `api/mcp.py`.
4. Rollback is simply leaving `AI_PLANNER_ENABLED` unset/false, or reverting
the changed files; no other capability depends on this one.
5. Enabling in a real environment requires setting `AI_PLANNER_ENABLED=true`,
`AI_PLANNER_PROVIDER` (`anthropic` or `openai`), and the corresponding
provider's API key in the process environment (already-existing SDK
convention, not a new config surface this change introduces).
## Open Questions
- Whether a `wait`/no-op tool is actually needed in practice, or whether the
self-correcting-next-turn behavior of a spurious action is good enough
indefinitely — deferred to real usage observation, not decided here.
- Whether `AI_PLANNER_TIMEOUT_SECONDS`'s default (30s) is well-tuned against
real provider latency for image-bearing requests — needs tuning against
real usage data once this is enabled somewhere with real traffic; not
fixed by this design.
- Whether provider/model selection should eventually move from
environment-variable/deployment-time to a per-task or per-request choice —
leaning toward "not until a concrete need appears" (YAGNI), consistent
with `semantic-scene-runtime`'s D4 reasoning for its own model choice.