feat(host-agent): persist UI-tree evidence and add overlay/action visualization
Tests / Test passed: 863
Tests / Test passed: 863
Fixes issue 3: the host-agent console showed OCR results but never real UI-tree data, because _ui_tree_nodes() checked for a get_ui_tree/ui_tree tool action that has never existed anywhere in the codebase. - storage/timeline.py: add a ui_tree_results field to TimelineRecord and Timeline.append(), mirroring the existing ocr_results field. - runtime/task.py: _append_timeline() now extracts scene.elements with source == "ui" into ui_tree_results (scene_builder.build_scene() already preserved these; they were just never persisted). - host_agent/web/app.py: _ui_tree_nodes() reads the new field directly instead of the dead tool-action check. New _overlay_payload() exposes each step's scene dimensions and fused element list for client-side rendering. - task_detail.html: adds a toggle to overlay OCR (orange) and UI-tree (blue) bounding boxes on the before-action screenshot, plus a visual marker for the actually executed action (tap circle, or an animated swipe path) using an SVG viewBox so no manual coordinate-scaling JS is needed. Legacy/incomplete records degrade to no overlay, never an error. Also corrects openspec/specs/runtime-task-evidence and host-agent-console-task-pages, which had encoded the same nonexistent-tool assumption, via the new host-agent-console-visual-evidence change. 600 tests passing; ruff/compileall/openspec validate all clean.
This commit is contained in:
@@ -0,0 +1,72 @@
|
||||
## Why
|
||||
|
||||
The Host Agent console's task-detail page only ever shows OCR observations,
|
||||
never UI-tree elements, even though every planning step fuses both into the
|
||||
step's `Scene`. The page's UI-tree rendering has been dead code since it was
|
||||
written: `_ui_tree_nodes()` only returns data for a step whose tool call is
|
||||
named `get_ui_tree`/`ui_tree` — an action that has never existed in
|
||||
`ALL_TOOL_SPECS` (the planner only ever calls `tap`/`swipe`/`input_text`/
|
||||
`launch_app`/`terminate_app`/`finish_task`). The `runtime-task-evidence` spec
|
||||
itself was written against this same wrong premise.
|
||||
|
||||
Separately, operators reviewing a task's evidence today can only look at
|
||||
raw before/after screenshots side by side. They asked for two additional
|
||||
ways to see the AI's actual perception and effect at a glance: (1) an
|
||||
on/off toggle that overlays the OCR/UI-tree bounding boxes directly on the
|
||||
screenshot, and (2) a visualization of what the executed action actually
|
||||
did — where a `tap` landed, or where a `swipe` started and ended — so a
|
||||
wrong tap/swipe target is visible without cross-referencing raw JSON
|
||||
coordinates against the image by hand.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Fix UI-tree evidence capture: `TaskRunner._append_timeline()` extracts the
|
||||
fused Scene's `source == "ui"` elements (already computed by
|
||||
`scene_builder.build_scene()` for every step) into a new
|
||||
`ui_tree_results` field on `TimelineRecord`/`Timeline.append()`, mirroring
|
||||
the existing `ocr_results` field.
|
||||
- Replace `host_agent/web/app.py::_ui_tree_nodes()`'s dead tool-name check
|
||||
with a direct read of the new `ui_tree_results` field, so the
|
||||
already-existing (previously unreachable) list markup in
|
||||
`task_detail.html` renders real data.
|
||||
- Add an operator-facing toggle on the task-detail page that overlays OCR
|
||||
and UI-tree bounding boxes (with label/text and source) directly on the
|
||||
**before**-screenshot — the frame the persisted `Scene` and its bounds
|
||||
actually describe. Purely client-side (HTML/CSS/JS): boxes are
|
||||
positioned from the bounds/coordinates already serialized into the
|
||||
rendered page; no new backend endpoint or persisted data.
|
||||
- Add an action-effect visualization on the same before-screenshot: a
|
||||
marker at the `tap` coordinates, or an animated start→end path for
|
||||
`swipe`, derived from the step's already-persisted `tool_call.args`.
|
||||
Non-spatial actions (`input_text`, `launch_app`, `terminate_app`,
|
||||
`finish_task`) render no marker.
|
||||
- Correct the `runtime-task-evidence` spec requirement that assumed a
|
||||
`get_ui_tree`/`ui_tree` tool call triggers UI-tree capture; it is
|
||||
rewritten to describe capture from every step's fused Scene.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `runtime-task-evidence`: UI-tree inspection requirement rewritten to
|
||||
source elements from the per-step fused Scene (`source == "ui"`
|
||||
elements) instead of a nonexistent `get_ui_tree`/`ui_tree` tool call.
|
||||
- `host-agent-console-task-pages`: task-detail page renders real UI-tree
|
||||
elements as a list (previously always empty); adds the OCR/UI-tree
|
||||
bounding-box overlay toggle and the tap/swipe action-effect
|
||||
visualization, both on the before-screenshot.
|
||||
|
||||
## Impact
|
||||
|
||||
- `runtime/task.py` — `TaskRunner._append_timeline()`
|
||||
- `storage/timeline.py` — `TimelineRecord`, `Timeline.append()`
|
||||
- `apps/device-host-agent/host_agent/web/app.py` — `_ui_tree_nodes()`,
|
||||
`_timeline_step_context()`
|
||||
- `apps/device-host-agent/host_agent/web/templates/task_detail.html` —
|
||||
overlay toggle control, overlay layer markup, action-effect markup
|
||||
- `apps/device-host-agent/host_agent/web/static/` — new client-side JS for
|
||||
overlay positioning/scaling and the action-effect animation (no new
|
||||
backend endpoint; all data is already inline in the rendered page)
|
||||
- No changes to `perception/scene_builder.py`'s fusion logic — it already
|
||||
preserves every UI-tree element with `source == "ui"` in `Scene.elements`
|
||||
- No new external dependencies
|
||||
Reference in New Issue
Block a user