Files
agentic-mobile-control/openspec/changes/host-agent-console-visual-evidence/proposal.md
T
q792602257 8162509158
Tests / Test passed: 863
feat(host-agent): persist UI-tree evidence and add overlay/action visualization
Fixes issue 3: the host-agent console showed OCR results but never real
UI-tree data, because _ui_tree_nodes() checked for a get_ui_tree/ui_tree
tool action that has never existed anywhere in the codebase.

- storage/timeline.py: add a ui_tree_results field to TimelineRecord and
  Timeline.append(), mirroring the existing ocr_results field.
- runtime/task.py: _append_timeline() now extracts scene.elements with
  source == "ui" into ui_tree_results (scene_builder.build_scene() already
  preserved these; they were just never persisted).
- host_agent/web/app.py: _ui_tree_nodes() reads the new field directly
  instead of the dead tool-action check. New _overlay_payload() exposes
  each step's scene dimensions and fused element list for client-side
  rendering.
- task_detail.html: adds a toggle to overlay OCR (orange) and UI-tree
  (blue) bounding boxes on the before-action screenshot, plus a visual
  marker for the actually executed action (tap circle, or an animated
  swipe path) using an SVG viewBox so no manual coordinate-scaling JS is
  needed. Legacy/incomplete records degrade to no overlay, never an error.

Also corrects openspec/specs/runtime-task-evidence and
host-agent-console-task-pages, which had encoded the same nonexistent-tool
assumption, via the new host-agent-console-visual-evidence change.

600 tests passing; ruff/compileall/openspec validate all clean.
2026-07-15 14:39:28 +08:00

73 lines
3.8 KiB
Markdown

## Why
The Host Agent console's task-detail page only ever shows OCR observations,
never UI-tree elements, even though every planning step fuses both into the
step's `Scene`. The page's UI-tree rendering has been dead code since it was
written: `_ui_tree_nodes()` only returns data for a step whose tool call is
named `get_ui_tree`/`ui_tree` — an action that has never existed in
`ALL_TOOL_SPECS` (the planner only ever calls `tap`/`swipe`/`input_text`/
`launch_app`/`terminate_app`/`finish_task`). The `runtime-task-evidence` spec
itself was written against this same wrong premise.
Separately, operators reviewing a task's evidence today can only look at
raw before/after screenshots side by side. They asked for two additional
ways to see the AI's actual perception and effect at a glance: (1) an
on/off toggle that overlays the OCR/UI-tree bounding boxes directly on the
screenshot, and (2) a visualization of what the executed action actually
did — where a `tap` landed, or where a `swipe` started and ended — so a
wrong tap/swipe target is visible without cross-referencing raw JSON
coordinates against the image by hand.
## What Changes
- Fix UI-tree evidence capture: `TaskRunner._append_timeline()` extracts the
fused Scene's `source == "ui"` elements (already computed by
`scene_builder.build_scene()` for every step) into a new
`ui_tree_results` field on `TimelineRecord`/`Timeline.append()`, mirroring
the existing `ocr_results` field.
- Replace `host_agent/web/app.py::_ui_tree_nodes()`'s dead tool-name check
with a direct read of the new `ui_tree_results` field, so the
already-existing (previously unreachable) list markup in
`task_detail.html` renders real data.
- Add an operator-facing toggle on the task-detail page that overlays OCR
and UI-tree bounding boxes (with label/text and source) directly on the
**before**-screenshot — the frame the persisted `Scene` and its bounds
actually describe. Purely client-side (HTML/CSS/JS): boxes are
positioned from the bounds/coordinates already serialized into the
rendered page; no new backend endpoint or persisted data.
- Add an action-effect visualization on the same before-screenshot: a
marker at the `tap` coordinates, or an animated start→end path for
`swipe`, derived from the step's already-persisted `tool_call.args`.
Non-spatial actions (`input_text`, `launch_app`, `terminate_app`,
`finish_task`) render no marker.
- Correct the `runtime-task-evidence` spec requirement that assumed a
`get_ui_tree`/`ui_tree` tool call triggers UI-tree capture; it is
rewritten to describe capture from every step's fused Scene.
## Capabilities
### Modified Capabilities
- `runtime-task-evidence`: UI-tree inspection requirement rewritten to
source elements from the per-step fused Scene (`source == "ui"`
elements) instead of a nonexistent `get_ui_tree`/`ui_tree` tool call.
- `host-agent-console-task-pages`: task-detail page renders real UI-tree
elements as a list (previously always empty); adds the OCR/UI-tree
bounding-box overlay toggle and the tap/swipe action-effect
visualization, both on the before-screenshot.
## Impact
- `runtime/task.py``TaskRunner._append_timeline()`
- `storage/timeline.py``TimelineRecord`, `Timeline.append()`
- `apps/device-host-agent/host_agent/web/app.py``_ui_tree_nodes()`,
`_timeline_step_context()`
- `apps/device-host-agent/host_agent/web/templates/task_detail.html`
overlay toggle control, overlay layer markup, action-effect markup
- `apps/device-host-agent/host_agent/web/static/` — new client-side JS for
overlay positioning/scaling and the action-effect animation (no new
backend endpoint; all data is already inline in the rendered page)
- No changes to `perception/scene_builder.py`'s fusion logic — it already
preserves every UI-tree element with `source == "ui"` in `Scene.elements`
- No new external dependencies