Files
q792602257 8162509158
Tests / Test passed: 863
feat(host-agent): persist UI-tree evidence and add overlay/action visualization
Fixes issue 3: the host-agent console showed OCR results but never real
UI-tree data, because _ui_tree_nodes() checked for a get_ui_tree/ui_tree
tool action that has never existed anywhere in the codebase.

- storage/timeline.py: add a ui_tree_results field to TimelineRecord and
  Timeline.append(), mirroring the existing ocr_results field.
- runtime/task.py: _append_timeline() now extracts scene.elements with
  source == "ui" into ui_tree_results (scene_builder.build_scene() already
  preserved these; they were just never persisted).
- host_agent/web/app.py: _ui_tree_nodes() reads the new field directly
  instead of the dead tool-action check. New _overlay_payload() exposes
  each step's scene dimensions and fused element list for client-side
  rendering.
- task_detail.html: adds a toggle to overlay OCR (orange) and UI-tree
  (blue) bounding boxes on the before-action screenshot, plus a visual
  marker for the actually executed action (tap circle, or an animated
  swipe path) using an SVG viewBox so no manual coordinate-scaling JS is
  needed. Legacy/incomplete records degrade to no overlay, never an error.

Also corrects openspec/specs/runtime-task-evidence and
host-agent-console-task-pages, which had encoded the same nonexistent-tool
assumption, via the new host-agent-console-visual-evidence change.

600 tests passing; ruff/compileall/openspec validate all clean.
2026-07-15 14:39:28 +08:00

3.8 KiB

Why

The Host Agent console's task-detail page only ever shows OCR observations, never UI-tree elements, even though every planning step fuses both into the step's Scene. The page's UI-tree rendering has been dead code since it was written: _ui_tree_nodes() only returns data for a step whose tool call is named get_ui_tree/ui_tree — an action that has never existed in ALL_TOOL_SPECS (the planner only ever calls tap/swipe/input_text/ launch_app/terminate_app/finish_task). The runtime-task-evidence spec itself was written against this same wrong premise.

Separately, operators reviewing a task's evidence today can only look at raw before/after screenshots side by side. They asked for two additional ways to see the AI's actual perception and effect at a glance: (1) an on/off toggle that overlays the OCR/UI-tree bounding boxes directly on the screenshot, and (2) a visualization of what the executed action actually did — where a tap landed, or where a swipe started and ended — so a wrong tap/swipe target is visible without cross-referencing raw JSON coordinates against the image by hand.

What Changes

  • Fix UI-tree evidence capture: TaskRunner._append_timeline() extracts the fused Scene's source == "ui" elements (already computed by scene_builder.build_scene() for every step) into a new ui_tree_results field on TimelineRecord/Timeline.append(), mirroring the existing ocr_results field.
  • Replace host_agent/web/app.py::_ui_tree_nodes()'s dead tool-name check with a direct read of the new ui_tree_results field, so the already-existing (previously unreachable) list markup in task_detail.html renders real data.
  • Add an operator-facing toggle on the task-detail page that overlays OCR and UI-tree bounding boxes (with label/text and source) directly on the before-screenshot — the frame the persisted Scene and its bounds actually describe. Purely client-side (HTML/CSS/JS): boxes are positioned from the bounds/coordinates already serialized into the rendered page; no new backend endpoint or persisted data.
  • Add an action-effect visualization on the same before-screenshot: a marker at the tap coordinates, or an animated start→end path for swipe, derived from the step's already-persisted tool_call.args. Non-spatial actions (input_text, launch_app, terminate_app, finish_task) render no marker.
  • Correct the runtime-task-evidence spec requirement that assumed a get_ui_tree/ui_tree tool call triggers UI-tree capture; it is rewritten to describe capture from every step's fused Scene.

Capabilities

Modified Capabilities

  • runtime-task-evidence: UI-tree inspection requirement rewritten to source elements from the per-step fused Scene (source == "ui" elements) instead of a nonexistent get_ui_tree/ui_tree tool call.
  • host-agent-console-task-pages: task-detail page renders real UI-tree elements as a list (previously always empty); adds the OCR/UI-tree bounding-box overlay toggle and the tap/swipe action-effect visualization, both on the before-screenshot.

Impact

  • runtime/task.pyTaskRunner._append_timeline()
  • storage/timeline.pyTimelineRecord, Timeline.append()
  • apps/device-host-agent/host_agent/web/app.py_ui_tree_nodes(), _timeline_step_context()
  • apps/device-host-agent/host_agent/web/templates/task_detail.html — overlay toggle control, overlay layer markup, action-effect markup
  • apps/device-host-agent/host_agent/web/static/ — new client-side JS for overlay positioning/scaling and the action-effect animation (no new backend endpoint; all data is already inline in the rendered page)
  • No changes to perception/scene_builder.py's fusion logic — it already preserves every UI-tree element with source == "ui" in Scene.elements
  • No new external dependencies