Files
agentic-mobile-control/openspec/changes/archive/2026-07-14-ai-planner-runtime/proposal.md
T
q792602257andClaude Opus 4.6 c01dd4c6b2 chore(openspec): archive ai-planner-runtime
Change is complete (22/22 tasks) and its delta spec has been synced
into a new main spec openspec/specs/agent-runtime/spec.md.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-14 23:28:17 +08:00

107 lines
6.1 KiB
Markdown

## Why
`runtime/planner.py::Planner` — the Planner half of Milestone 3 (Agent
Runtime)'s `agent-runtime` capability, implemented by `apex-agent-mvp` and
still the only Planner in the codebase — is a stub: it always returns exactly
one hardcoded `describe_screen` step on a task's first call, then `[]` on
every call after any step has executed, regardless of the goal or the current
`Scene`. `TaskRunner`'s Observe→Plan→Act→Observe loop, `Executor`'s
retry/backoff, `TaskContext`'s per-task memory, and the already-implemented
`semantic-scene` and `world-model` capabilities are all real and wired
end-to-end — every later milestone (Task Memory, Semantic Scene, World Model,
Skill Learning, Workflow Orchestration, Multi-Agent Runtime) has been built on
top of, or alongside, this stub without ever replacing it. Nothing in the
runtime actually decides what to do. This change gives the runtime its first
real decision-maker: an LLM-driven Planner that uses native tool/function
calling to choose exactly one grounded action per turn, completing what
Milestone 3 always intended the Planner role to be.
## What Changes
- Add `runtime/ai_planner.py::AIPlanner`, a `Planner` implementation that
calls an LLM with native tool/function calling once per `plan()`
invocation, translating the model's single chosen tool call into 0 or 1
`PlannedStep`. Single-step, ReAct-style: `TaskRunner.run()` already
re-observes and re-plans every iteration, so `AIPlanner` never attempts
multi-step lookahead.
- Add a **pluggable dual-provider** LLM abstraction
(`runtime/tool_calling_client.py`): both Anthropic native tool use and
OpenAI function calling are supported, selected via configuration
(`AI_PLANNER_PROVIDER`), each forced to return exactly one tool call per
turn so a response always resolves to a single, unambiguous decision.
- Add a fixed, six-tool tool-face (`runtime/tool_specs.py`): `tap`, `swipe`,
`input_text`, `launch_app`, `terminate_app` for action, plus an explicit
`finish_task(success, reason)` control tool the model calls to end the task
(on success or failure) instead of relying on an ambiguous "no tool call"
signal.
- **Amend the Perception Boundary invariant** (`docs/CONSTITUTION.md`) with
one narrow, explicit exception: the AI `Planner` (only that Planner) may
additionally receive the current step's raw screenshot bytes alongside
`Scene`, to support vision-grounded action grounding (precise tap/swipe
coordinates). No other layer or LLM consumer gains access to raw screenshot
bytes; `Scene` remains the only perception artifact everywhere else.
- Add configuration (`runtime/planner_config.py`: `AI_PLANNER_ENABLED`,
`AI_PLANNER_PROVIDER`, `AI_PLANNER_MODEL`, `AI_PLANNER_TIMEOUT_SECONDS`)
defaulting `AI_PLANNER_ENABLED=False`, matching the existing enable/disable
convention used by every other LLM-backed capability (`semantic-scene`,
`skill-learning`). `runtime/task.py::TaskRunner` builds an `AIPlanner` by
default only when enabled; otherwise its existing stub `Planner` behavior
is unchanged.
- Fix a latent correctness gap in `TaskRunner.run()`'s loop, exposed by
giving the Planner a real chance to fail: observe/plan exceptions are now
caught per iteration and turned into a `status="failed"` task with a
`failure_reason`, instead of propagating uncaught (previously harmless only
because the stub Planner never raised).
- **BREAKING**: none. Default-disabled; when disabled, `TaskRunner`'s
constructed `Planner` and control flow are unchanged from before this
change.
## Capabilities
### New Capabilities
(none — this change fulfills the Planner role already scoped by Milestone
3's `agent-runtime` capability; see Impact for why no new capability is
declared)
### Modified Capabilities
- `agent-runtime`: the Planner requirement ("given a goal and the current
Scene, produces steps") gains a real, LLM-backed implementation with
native tool calling, a defined single-action-per-turn contract, an
explicit task-completion/failure signal (`finish_task`), and — as a
Planner-only exception to the Perception Boundary — optional access to the
current step's screenshot. `agent-runtime`'s base spec was never archived
into `openspec/specs/` (a pre-existing gap from `apex-agent-mvp`, out of
scope for this change); this change's delta spec below uses `## ADDED
Requirements` against that as-yet-unarchived baseline, the same way
`semantic-scene-runtime` and `world-model-runtime` each layered their own
delta on top of it without attempting to backfill it (see `design.md`).
## Impact
- **New files**: `runtime/planner_config.py`, `runtime/tool_specs.py`,
`runtime/tool_calling_client.py`, `runtime/planner_prompts.py`,
`runtime/ai_planner.py`.
- **Modified**: `runtime/planner.py` (base `Planner.plan()` gains an optional
`screenshot` parameter, stub behavior unchanged), `runtime/task.py`
(`TaskRunner` builds an `AIPlanner` when enabled; generalized its existing
`_planner_accepts_world()` reflection helper to also conditionally inject
`screenshot`; wrapped the observe+plan step in try/except),
`docs/CONSTITUTION.md` (Perception Boundary amendment described above).
- **No change** to `api/rest.py``TaskRunner()`'s existing zero-argument
construction picks up the new behavior automatically once
`AI_PLANNER_ENABLED=true` is set in the environment; no new request
parameters, no MCP surface change (`api/mcp.py` is a separate, curated
tool-handler surface unaffected by this change).
- **No new dependencies**: `pyproject.toml` already declares both
`anthropic` and `openai` SDKs (the former added for
`semantic-scene-runtime`); this change is the first to actually construct
an OpenAI client.
- **Out of scope**: no `wait`/no-op tool (v1.1 candidate, see `design.md`);
no dynamic tool subset selection; no provider/model choice exposed as an
API request parameter; no automatic retry of failed LLM calls inside
`AIPlanner` (transient failures surface as a failed task; tool-execution
retry remains solely `Executor`'s concern, unchanged by this change); no
backfill of the missing `openspec/specs/agent-runtime/` base spec.