feat(agent-runtime): add LLM-driven AI Planner with dual-provider tool calling
Replaces the stub Planner's fixed describe_screen/[] behavior with a real decision-maker: AIPlanner uses native tool/function calling (Anthropic or OpenAI, pluggable via AI_PLANNER_PROVIDER) to select exactly one grounded action per turn, with an explicit finish_task(success, reason) tool for completion/failure instead of an ambiguous "no tool call" signal. Default disabled (AI_PLANNER_ENABLED=false) and additive; TaskRunner falls back to the existing stub Planner unchanged when disabled. Amends CONSTITUTION.md's Perception Boundary with one narrow exception: only the AI Planner may receive the current step's raw screenshot bytes alongside Scene, for vision-grounded coordinate grounding. Also fixes a latent gap in TaskRunner.run(): observe/plan exceptions are now caught per iteration and turned into a failed task with a failure_reason, instead of propagating uncaught. openspec change: ai-planner-runtime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -21,7 +21,13 @@ driver such as `WDADriver`.
|
||||
|
||||
## Perception Boundary
|
||||
|
||||
`Scene` is the only perception artifact the LLM sees. It is produced through
|
||||
`Scene` is the only perception artifact the LLM sees, with one narrow,
|
||||
explicit exception: the runtime-layer AI `Planner` (and only that Planner)
|
||||
may additionally receive the raw screenshot bytes for the current step,
|
||||
alongside `Scene`, to support vision-grounded decision-making. No other
|
||||
layer — `api`, `tools`, `perception`, `storage`, or any other LLM consumer —
|
||||
may receive raw screenshot bytes; every other perception consumer still
|
||||
receives `Scene` only. `Scene` itself is still produced exclusively through
|
||||
`PerceptionProvider`, not by direct calls to OCR, UI tree parsing, or
|
||||
`scene_builder` from runtime and API layers.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user