Files
2026-07-15 18:14:28 +08:00

14 KiB

Context

The AI Planner currently invokes one LLM call per step via AIPlanner.plan(), returning a PlannedStep with action/args. The ToolCallingClient extracts only the tool_use block from the response; any thinking or text blocks in the Anthropic response are silently discarded.

WorldModel._append_history() records one WorldEvent per step, storing the full scene_summary (SemanticScene or Scene). _history_summary() in ai_planner.py serialises these as full JSON for the next prompt, causing rapid token growth. The planner has no mechanism to reflect on whether a prior action achieved its intended effect.

Current state of the pipeline:

plan(scene, world) → user_prompt → LLM → [thinking?, text?, tool_use] → only tool_use kept
                                                                          ↓
WorldEvent(scene_summary=<full UI tree>, action, success) → history → next prompt (fat)

Target state:

plan(scene, world) → user_prompt → LLM → thinking (kept) + text=rationale (kept) + tool_use
                                                                          ↓
WorldEvent(rationale, thinking, current_page, action, success) → compact history → next prompt

Constraints:

  • runtime/ package must not import cloud or host_agent concerns (enforced by existing test).
  • CloudProxyToolCallingClient lives in host_agent/ and must remain opaque to this change — it forwards raw prompts to the Cloud API and receives tool_name/arguments back; thinking/text are not surfaced in that transport path.
  • Extended thinking requires Anthropic SDK and is Anthropic-only; must be opt-in and degrade gracefully when not configured.
  • WorldEvent schema change must be backward-compatible: existing stored events without rationale/thinking fields must remain readable.

Goals / Non-Goals

Goals:

  • Capture AI rationale (pre-tool text reflection) from every planning call where the model outputs one.
  • Capture extended thinking when AI_PLANNER_THINKING_BUDGET_TOKENS is configured.
  • Reduce per-step history token cost by replacing scene_summary with compact {page, rationale, action, success} in the prompt history format.
  • Prompt the AI to evaluate the previous step's outcome before selecting the next action (method C: same LLM call, no extra API round-trip).
  • Persist thinking and rationale in the Cloud-side planner_decision_log table alongside existing prompt/tool records.
  • Capture OpenAI reasoning_content when present (o-series models).

Non-Goals:

  • Not adding a dedicated reflection LLM call (explicit method B) — method C (pre-tool text block) achieves the intent within the existing call budget.
  • Not changing the scene_summary field in WorldEvent for stored/displayed task history in the Console — only the prompt-construction path (_history_summary()) is changed.
  • Not making thinking_budget_tokens configurable per-task at runtime (only via environment variable).

Decisions

D1: Method C — reflection within the same LLM call via pre-tool text block

Decision: Update PLANNER_SYSTEM_PROMPT to instruct the AI to output a short text block before the tool call assessing the previous step's outcome and stating the current step's intent. Extract this text block as text_output alongside the tool_use block.

Why over method B (explicit reflection call): Method B doubles LLM calls per step (and associated latency/cost). The AI already has the previous action and current scene available; requiring a text block costs one extra sentence in the response, not a round-trip.

Why over method A (implicit): Method A relies on the AI organically reflecting without prompting, which is unreliable. Explicit instruction in the system prompt makes it consistent.

Correction (post-implementation, found during manual smoke testing per task 8.5): The original assumption that tool_choice: {type: "any"} allows text blocks before tool_use was wrong. Per Anthropic's API behaviour, forced tool_choice (any or a specific tool) makes Claude skip any preceding text block entirely, and forced tool_choice is also incompatible with extended thinking. Under the original tool_choice: {type: "any"} call, text_output/thinking were therefore always None in practice — not merely "sometimes absent" as D1/Risks originally assumed. See D8 for the fix.

D8: tool_choice must be "auto" (with a forced retry fallback) to allow rationale/thinking capture

Decision: AnthropicToolCallingClient/OpenAIToolCallingClient now call the API with tool_choice: "auto" first (Anthropic: {"type": "auto", "disable_parallel_tool_use": True}; OpenAI: "auto"). Extended thinking (when thinking_budget_tokens is set) is only ever requested alongside "auto". If the model responds without any tool call at all, the client retries once with the original forced tool_choice (Anthropic "any", OpenAI "required") and no thinking parameter, guaranteeing a tool call is eventually returned. OpenAIToolCallingClient also now captures message.content as text_output (previously never extracted for any provider — OpenAI never had rationale capture at all, forced or not).

Why not just always force tool_choice with a "think first" instruction: Confirmed via Anthropic's official docs/SDK guidance that this combination structurally suppresses the text/thinking Claude would otherwise produce — no amount of prompting fixes it while tool_choice stays forced.

Why a fallback retry rather than failing the step: PLANNER_SYSTEM_PROMPT already instructs "you must then call exactly one tool", so tool_choice: "auto" calls overwhelmingly still return a tool_use block; the retry only guards the rare case where the model responds with pure text. Failing the task step outright on that rare case would regress reliability for a cosmetic (rationale`) improvement. The forced retry accepts losing rationale/thinking for that one step rather than losing task progress.

D2: Extend ToolCallDecision with thinking and text_output

Decision: Add thinking: str | None = None and text_output: str | None = None to ToolCallDecision. _decision_from_anthropic_response() iterates all content blocks, collecting the first thinking block's text and concatenating any text blocks. _decision_from_openai_response() extracts reasoning_content from the first choice's message when present.

Why not a separate data structure: ToolCallDecision is the single return type of decide(); adding optional fields keeps the interface minimal and backward-compatible for CloudProxyToolCallingClient which sets neither field.

D3: Extended thinking via thinking_budget_tokens in PlannerConfig

Decision: Add thinking_budget_tokens: int | None = None to PlannerConfig, loaded from AI_PLANNER_THINKING_BUDGET_TOKENS env var. When set and provider is anthropic, AnthropicToolCallingClient._create_message() adds thinking: {type: "enabled", budget_tokens: N} and the interleaved-thinking-2025-05-14 beta header. max_tokens must be ≥ budget_tokens + 1; the client enforces this by taking max(max_tokens, budget_tokens + 1).

Why opt-in via env var: Extended thinking increases latency and cost significantly. Most tasks don't need it. Keeping it off by default preserves existing behaviour exactly.

Why enforce max_tokens floor: The Anthropic API rejects requests where max_tokens ≤ budget_tokens; silent enforcement avoids a confusing runtime error.

D4: PlannedStep carries rationale and thinking

Decision: Add rationale: str | None = None and thinking: str | None = None to PlannedStep. AIPlanner.plan() populates these from decision.text_output and decision.thinking.

Why on PlannedStep: WorldModel._append_history() receives the PlannedStep; this is the natural carrier from the planner layer to the world model layer without adding a new coupling.

D5: WorldEvent gets rationale and thinking; scene_summary becomes optional

Decision: WorldEvent.scene_summary changes from SemanticScene | Scene to SemanticScene | Scene | None, defaulting to None. New fields rationale: str | None = None and thinking: str | None = None are added. _append_history() in WorldModel passes step.rationale and step.thinking and no longer requires a non-None scene_summary.

Why keep scene_summary as optional rather than removing it: Some callers (e.g., direct non-AI planner paths) still populate it. Keeping it optional is backward-compatible and avoids breaking the field in stored timelines.

Why current_page instead of scene_summary in prompt history: WorldState.current_page is a single string (e.g. "Settings > Account"), already maintained by WorldModel. Including it in the compact history record gives the AI verifiable page-level context to validate the previous step's navigation intent, at near-zero token cost.

D6: Compact history format in _history_summary()

Decision: _history_summary() in ai_planner.py switches from [event.to_dict() for event in world.history] to a compact list:

[{
    "page": event.page,
    "action": event.action,
    "rationale": event.rationale,
    "success": event.success,
} for event in world.history]

event.page is added as a field to WorldEvent (populated from WorldState.current_page at the time of recording).

Why not expose thinking in history: Thinking blocks are verbose internal reasoning, not summaries. Exposing them in history would re-introduce token bloat. Rationale (the AI's own compact 1–2 sentence reflection) is the right unit for history.

D7: planner_decision_log table extended with thinking and rationale

Decision: Add nullable thinking TEXT and rationale TEXT columns to planner_decision_log via a new Alembic migration. record_planner_decision() in cloud/internal_api/api.py and PlannerDecisionRecord in cloud/internal_api/models.py are extended to carry these fields. Existing rows without the columns read as NULL.

Why in the cloud decision log: The cloud-proxy path receives the full PlannerDecisionRecord including thinking/rationale from the Host Agent. Persisting them completes the audit trail for cloud-transport tasks.

Why separate columns rather than a JSON blob: The existing table uses discrete columns for system_prompt, user_prompt, tool_name, tool_arguments — consistency favours discrete columns. Both fields are optional (cloud-proxy path only; direct transport never produces cloud decision records).

D9: Device-action tool calls carry required purpose and expected outcome

Decision: Every device-action schema (tap, swipe, input_text, launch_app, and terminate_app) adds required non-empty purpose and expected_outcome string fields. The tool-call response parser removes these metadata fields from executable arguments and assigns them to ToolCallDecision; AIPlanner carries them on PlannedStep. The completion signal remains unchanged because it already has a required terminal reason.

The Cloud planner response returns rationale, thinking, purpose, and expected outcome to the Host. The Cloud decision log persists the four values as nullable, additive columns so historic rows and non-AI callers remain readable. WorldEvent retains the executable action arguments together with purpose and expected outcome; its compact history summary includes that action record. Timeline records and learned FlowStep instances retain the same arguments and metadata; flow embedding text includes available semantics so retrieval can use them.

Why structured tool arguments rather than rationale: the pre-tool text block is optional by API design and can be suppressed by a forced retry. Tool schemas are the provider-enforced structured-output boundary, so requiring purpose and expected outcome there makes them available for every submitted device action without relying on free-form rationale.

Why strip metadata before execution: device tool functions only accept their physical-action arguments. Keeping metadata off step.args preserves their API contracts while still making it available to execution history, verification, Cloud audit, and skill synthesis.

Risks / Trade-offs

  • AI may not always output a text block: even with tool_choice: "auto" (D8), the model is not guaranteed to prefix a text block before the tool call. When absent, text_output is None and rationale is None. History degrades gracefully to {page, rationale: null, action, success}.
  • tool_choice: "auto" occasionally yields no tool call at all: unlike forced tool_choice, "auto" permits the model to respond with text only and no tool call. D8's forced retry (no thinking, no rationale on that path) guards this case so a step never stalls; this trades away rationale/thinking for that single step, not overall reliability.
  • Extended thinking increases latency: budget_tokens directly adds to minimum response time. This is opt-in and accepted by the operator who enables it.
  • WorldEvent schema divergence from stored data: Existing WorldEvent instances in memory or serialised timelines lack rationale/thinking. The to_dict() method will emit null for these fields; downstream consumers should treat null as absent, not as a failure.
  • Cloud/Host deployment order: a new Host requires a Cloud API that returns the additive metadata fields to preserve them locally. The Cloud response models remain nullable so an old peer remains readable during a rolling deployment, but it cannot provide the new semantic records.

Migration Plan

  1. Apply Alembic migration for new planner_decision_log columns (additive, nullable, no data migration required).
  2. Deploy Cloud API with updated PlannerDecisionRecord and record_planner_decision().
  3. Deploy Host Agent with updated tool_calling_client, planner_config, ai_planner, planner_prompts, world/models, world/model.
  4. Rollback: revert Host Agent to prior version (no DB rollback needed — new columns simply go unused).

Open Questions

  • None blocking implementation.