14 KiB
Context
The AI Planner currently invokes one LLM call per step via AIPlanner.plan(), returning a PlannedStep with action/args. The ToolCallingClient extracts only the tool_use block from the response; any thinking or text blocks in the Anthropic response are silently discarded.
WorldModel._append_history() records one WorldEvent per step, storing the full scene_summary (SemanticScene or Scene). _history_summary() in ai_planner.py serialises these as full JSON for the next prompt, causing rapid token growth. The planner has no mechanism to reflect on whether a prior action achieved its intended effect.
Current state of the pipeline:
plan(scene, world) → user_prompt → LLM → [thinking?, text?, tool_use] → only tool_use kept
↓
WorldEvent(scene_summary=<full UI tree>, action, success) → history → next prompt (fat)
Target state:
plan(scene, world) → user_prompt → LLM → thinking (kept) + text=rationale (kept) + tool_use
↓
WorldEvent(rationale, thinking, current_page, action, success) → compact history → next prompt
Constraints:
runtime/package must not importcloudorhost_agentconcerns (enforced by existing test).CloudProxyToolCallingClientlives inhost_agent/and must remain opaque to this change — it forwards raw prompts to the Cloud API and receivestool_name/argumentsback; thinking/text are not surfaced in that transport path.- Extended thinking requires Anthropic SDK and is Anthropic-only; must be opt-in and degrade gracefully when not configured.
WorldEventschema change must be backward-compatible: existing stored events withoutrationale/thinkingfields must remain readable.
Goals / Non-Goals
Goals:
- Capture AI rationale (pre-tool text reflection) from every planning call where the model outputs one.
- Capture extended thinking when
AI_PLANNER_THINKING_BUDGET_TOKENSis configured. - Reduce per-step history token cost by replacing
scene_summarywith compact{page, rationale, action, success}in the prompt history format. - Prompt the AI to evaluate the previous step's outcome before selecting the next action (method C: same LLM call, no extra API round-trip).
- Persist
thinkingandrationalein the Cloud-sideplanner_decision_logtable alongside existing prompt/tool records. - Capture OpenAI
reasoning_contentwhen present (o-series models).
Non-Goals:
- Not adding a dedicated reflection LLM call (explicit method B) — method C (pre-tool text block) achieves the intent within the existing call budget.
- Not changing the
scene_summaryfield inWorldEventfor stored/displayed task history in the Console — only the prompt-construction path (_history_summary()) is changed. - Not making
thinking_budget_tokensconfigurable per-task at runtime (only via environment variable).
Decisions
D1: Method C — reflection within the same LLM call via pre-tool text block
Decision: Update PLANNER_SYSTEM_PROMPT to instruct the AI to output a short text block before the tool call assessing the previous step's outcome and stating the current step's intent. Extract this text block as text_output alongside the tool_use block.
Why over method B (explicit reflection call): Method B doubles LLM calls per step (and associated latency/cost). The AI already has the previous action and current scene available; requiring a text block costs one extra sentence in the response, not a round-trip.
Why over method A (implicit): Method A relies on the AI organically reflecting without prompting, which is unreliable. Explicit instruction in the system prompt makes it consistent.
Correction (post-implementation, found during manual smoke testing per task 8.5): The original assumption that tool_choice: {type: "any"} allows text blocks before tool_use was wrong. Per Anthropic's API behaviour, forced tool_choice (any or a specific tool) makes Claude skip any preceding text block entirely, and forced tool_choice is also incompatible with extended thinking. Under the original tool_choice: {type: "any"} call, text_output/thinking were therefore always None in practice — not merely "sometimes absent" as D1/Risks originally assumed. See D8 for the fix.
D8: tool_choice must be "auto" (with a forced retry fallback) to allow rationale/thinking capture
Decision: AnthropicToolCallingClient/OpenAIToolCallingClient now call the API with tool_choice: "auto" first (Anthropic: {"type": "auto", "disable_parallel_tool_use": True}; OpenAI: "auto"). Extended thinking (when thinking_budget_tokens is set) is only ever requested alongside "auto". If the model responds without any tool call at all, the client retries once with the original forced tool_choice (Anthropic "any", OpenAI "required") and no thinking parameter, guaranteeing a tool call is eventually returned. OpenAIToolCallingClient also now captures message.content as text_output (previously never extracted for any provider — OpenAI never had rationale capture at all, forced or not).
Why not just always force tool_choice with a "think first" instruction: Confirmed via Anthropic's official docs/SDK guidance that this combination structurally suppresses the text/thinking Claude would otherwise produce — no amount of prompting fixes it while tool_choice stays forced.
Why a fallback retry rather than failing the step: PLANNER_SYSTEM_PROMPT already instructs "you must then call exactly one tool", so tool_choice: "auto" calls overwhelmingly still return a tool_use block; the retry only guards the rare case where the model responds with pure text. Failing the task step outright on that rare case would regress reliability for a cosmetic (rationale`) improvement. The forced retry accepts losing rationale/thinking for that one step rather than losing task progress.
D2: Extend ToolCallDecision with thinking and text_output
Decision: Add thinking: str | None = None and text_output: str | None = None to ToolCallDecision. _decision_from_anthropic_response() iterates all content blocks, collecting the first thinking block's text and concatenating any text blocks. _decision_from_openai_response() extracts reasoning_content from the first choice's message when present.
Why not a separate data structure: ToolCallDecision is the single return type of decide(); adding optional fields keeps the interface minimal and backward-compatible for CloudProxyToolCallingClient which sets neither field.
D3: Extended thinking via thinking_budget_tokens in PlannerConfig
Decision: Add thinking_budget_tokens: int | None = None to PlannerConfig, loaded from AI_PLANNER_THINKING_BUDGET_TOKENS env var. When set and provider is anthropic, AnthropicToolCallingClient._create_message() adds thinking: {type: "enabled", budget_tokens: N} and the interleaved-thinking-2025-05-14 beta header. max_tokens must be ≥ budget_tokens + 1; the client enforces this by taking max(max_tokens, budget_tokens + 1).
Why opt-in via env var: Extended thinking increases latency and cost significantly. Most tasks don't need it. Keeping it off by default preserves existing behaviour exactly.
Why enforce max_tokens floor: The Anthropic API rejects requests where max_tokens ≤ budget_tokens; silent enforcement avoids a confusing runtime error.
D4: PlannedStep carries rationale and thinking
Decision: Add rationale: str | None = None and thinking: str | None = None to PlannedStep. AIPlanner.plan() populates these from decision.text_output and decision.thinking.
Why on PlannedStep: WorldModel._append_history() receives the PlannedStep; this is the natural carrier from the planner layer to the world model layer without adding a new coupling.
D5: WorldEvent gets rationale and thinking; scene_summary becomes optional
Decision: WorldEvent.scene_summary changes from SemanticScene | Scene to SemanticScene | Scene | None, defaulting to None. New fields rationale: str | None = None and thinking: str | None = None are added. _append_history() in WorldModel passes step.rationale and step.thinking and no longer requires a non-None scene_summary.
Why keep scene_summary as optional rather than removing it: Some callers (e.g., direct non-AI planner paths) still populate it. Keeping it optional is backward-compatible and avoids breaking the field in stored timelines.
Why current_page instead of scene_summary in prompt history: WorldState.current_page is a single string (e.g. "Settings > Account"), already maintained by WorldModel. Including it in the compact history record gives the AI verifiable page-level context to validate the previous step's navigation intent, at near-zero token cost.
D6: Compact history format in _history_summary()
Decision: _history_summary() in ai_planner.py switches from [event.to_dict() for event in world.history] to a compact list:
[{
"page": event.page,
"action": event.action,
"rationale": event.rationale,
"success": event.success,
} for event in world.history]
event.page is added as a field to WorldEvent (populated from WorldState.current_page at the time of recording).
Why not expose thinking in history: Thinking blocks are verbose internal reasoning, not summaries. Exposing them in history would re-introduce token bloat. Rationale (the AI's own compact 1–2 sentence reflection) is the right unit for history.
D7: planner_decision_log table extended with thinking and rationale
Decision: Add nullable thinking TEXT and rationale TEXT columns to planner_decision_log via a new Alembic migration. record_planner_decision() in cloud/internal_api/api.py and PlannerDecisionRecord in cloud/internal_api/models.py are extended to carry these fields. Existing rows without the columns read as NULL.
Why in the cloud decision log: The cloud-proxy path receives the full PlannerDecisionRecord including thinking/rationale from the Host Agent. Persisting them completes the audit trail for cloud-transport tasks.
Why separate columns rather than a JSON blob: The existing table uses discrete columns for system_prompt, user_prompt, tool_name, tool_arguments — consistency favours discrete columns. Both fields are optional (cloud-proxy path only; direct transport never produces cloud decision records).
D9: Device-action tool calls carry required purpose and expected outcome
Decision: Every device-action schema (tap, swipe, input_text,
launch_app, and terminate_app) adds required non-empty purpose and
expected_outcome string fields. The tool-call response parser removes these
metadata fields from executable arguments and assigns them to
ToolCallDecision; AIPlanner carries them on PlannedStep. The completion
signal remains unchanged because it already has a required terminal reason.
The Cloud planner response returns rationale, thinking, purpose, and expected
outcome to the Host. The Cloud decision log persists the four values as
nullable, additive columns so historic rows and non-AI callers remain
readable. WorldEvent retains the executable action arguments together with
purpose and expected outcome; its compact history summary includes that action
record. Timeline records and learned FlowStep instances retain the same
arguments and metadata; flow embedding text includes available semantics so
retrieval can use them.
Why structured tool arguments rather than rationale: the pre-tool text block is optional by API design and can be suppressed by a forced retry. Tool schemas are the provider-enforced structured-output boundary, so requiring purpose and expected outcome there makes them available for every submitted device action without relying on free-form rationale.
Why strip metadata before execution: device tool functions only accept
their physical-action arguments. Keeping metadata off step.args preserves
their API contracts while still making it available to execution history,
verification, Cloud audit, and skill synthesis.
Risks / Trade-offs
- AI may not always output a text block: even with
tool_choice: "auto"(D8), the model is not guaranteed to prefix a text block before the tool call. When absent,text_outputisNoneandrationaleisNone. History degrades gracefully to{page, rationale: null, action, success}. tool_choice: "auto"occasionally yields no tool call at all: unlike forcedtool_choice,"auto"permits the model to respond with text only and no tool call. D8's forced retry (no thinking, no rationale on that path) guards this case so a step never stalls; this trades away rationale/thinking for that single step, not overall reliability.- Extended thinking increases latency:
budget_tokensdirectly adds to minimum response time. This is opt-in and accepted by the operator who enables it. WorldEventschema divergence from stored data: ExistingWorldEventinstances in memory or serialised timelines lackrationale/thinking. Theto_dict()method will emitnullfor these fields; downstream consumers should treatnullas absent, not as a failure.- Cloud/Host deployment order: a new Host requires a Cloud API that returns the additive metadata fields to preserve them locally. The Cloud response models remain nullable so an old peer remains readable during a rolling deployment, but it cannot provide the new semantic records.
Migration Plan
- Apply Alembic migration for new
planner_decision_logcolumns (additive, nullable, no data migration required). - Deploy Cloud API with updated
PlannerDecisionRecordandrecord_planner_decision(). - Deploy Host Agent with updated
tool_calling_client,planner_config,ai_planner,planner_prompts,world/models,world/model. - Rollback: revert Host Agent to prior version (no DB rollback needed — new columns simply go unused).
Open Questions
- None blocking implementation.