Files
agentic-mobile-control/openspec/changes/host-agent-console-visual-evidence/design.md
T
q792602257 8162509158
Tests / Test passed: 863
feat(host-agent): persist UI-tree evidence and add overlay/action visualization
Fixes issue 3: the host-agent console showed OCR results but never real
UI-tree data, because _ui_tree_nodes() checked for a get_ui_tree/ui_tree
tool action that has never existed anywhere in the codebase.

- storage/timeline.py: add a ui_tree_results field to TimelineRecord and
  Timeline.append(), mirroring the existing ocr_results field.
- runtime/task.py: _append_timeline() now extracts scene.elements with
  source == "ui" into ui_tree_results (scene_builder.build_scene() already
  preserved these; they were just never persisted).
- host_agent/web/app.py: _ui_tree_nodes() reads the new field directly
  instead of the dead tool-action check. New _overlay_payload() exposes
  each step's scene dimensions and fused element list for client-side
  rendering.
- task_detail.html: adds a toggle to overlay OCR (orange) and UI-tree
  (blue) bounding boxes on the before-action screenshot, plus a visual
  marker for the actually executed action (tap circle, or an animated
  swipe path) using an SVG viewBox so no manual coordinate-scaling JS is
  needed. Legacy/incomplete records degrade to no overlay, never an error.

Also corrects openspec/specs/runtime-task-evidence and
host-agent-console-task-pages, which had encoded the same nonexistent-tool
assumption, via the new host-agent-console-visual-evidence change.

600 tests passing; ruff/compileall/openspec validate all clean.
2026-07-15 14:39:28 +08:00

178 lines
9.2 KiB
Markdown

## Context
Every planning step already computes a fused `Scene` via
`perception/scene_builder.py::build_scene()`. Confirmed by reading the
fusion logic: every UI-tree element passed in as `ui_elements` survives
into `Scene.elements` with `source == "ui"` unchanged (whether or not it
matched an OCR box — matching only overwrites `text`/`confidence`, never
`source` or drops the element). Only OCR elements that never matched a
UI-tree element are additionally kept, separately, in `Scene.ocr_elements`.
`runtime/task.py::_append_timeline()` today only ever extracts
`source == "ocr"` elements (via `scene.ocr_results_to_dict()`, with a
fallback scan of `scene.elements`) into the Timeline. No equivalent
extraction of `source == "ui"` elements exists, and `TimelineRecord` has no
field to hold them even if extracted.
`host_agent/web/app.py::_ui_tree_nodes()` was written against a different,
incorrect assumption: that UI-tree data only exists when a step's tool call
is `get_ui_tree`/`ui_tree`. That action has never existed in
`runtime/tool_specs.py::ALL_TOOL_SPECS` (`tap`, `swipe`, `input_text`,
`launch_app`, `terminate_app`, `finish_task` only) — so `_ui_tree_nodes()`
returns `[]` for every real task. `task_detail.html` already has working
Jinja2 markup for `step.ui_tree_nodes` (a `<details>` block, same shape as
the OCR one), so once the data is wired through it needs no template
rewrite.
`tap`/`swipe` tool arguments (`runtime/tool_specs.py`) are specified in
"Scene pixel coordinates" — the same coordinate space as
`SceneElement.bounds` and `Scene.width`/`Scene.height`. This is what makes
drawing both the perception boxes and the action-effect marker from the
same coordinate space consistent.
## Goals / Non-Goals
**Goals:**
- Make UI-tree elements flow from the fused `Scene` into the Timeline and
render as a list on the task-detail page, replacing dead code.
- Let an operator toggle an overlay of OCR/UI-tree bounding boxes directly
on the screenshot that scene data was actually captured from.
- Visualize the actual spatial effect of `tap`/`swipe` actions on that same
screenshot.
- Do this without any new backend endpoint or additional persisted fields
beyond one new Timeline field (`ui_tree_results`), since bounds/text/tool
args are already computed and already serialized to the page today.
**Non-Goals:**
- Not changing `perception/scene_builder.py`'s fusion algorithm.
- Not adding overlay/animation to the after-screenshot — the persisted
`Scene` and its bounds describe the state the action was planned
against (the before-screenshot), not the resulting state. Overlaying
boxes on the after-image would misleadingly imply they describe
post-action element positions.
- Not persisting rendered overlay images; overlay/animation are computed
client-side from data already in the page.
- Not adding overlay/animation for non-spatial actions (`input_text`,
`launch_app`, `terminate_app`, `finish_task`).
## Decisions
### D1: UI-tree evidence sourced from `Scene.elements` where `source == "ui"`, not a new Scene field
**Decision**: `_append_timeline()` computes
`ui_tree_results = [element.to_dict() for element in scene.elements if element.source == "ui"]`
and passes it to `Timeline.append(..., ui_tree_results=ui_tree_results)`.
`TimelineRecord` gains `ui_tree_results: list[dict[str, Any]] = field(default_factory=list)`,
mirroring the existing `ocr_results` field exactly.
**Why not add a raw `ui_elements` field to `Scene`**: `scene_builder.build_scene()`
already preserves every UI-tree element in `Scene.elements` untouched aside
from an OCR-provided `text`/`confidence` merge; filtering by `source` at the
Timeline layer needs no change to the perception layer, and matches the
existing pattern for how `ocr_results` gets a fallback scan of
`scene.elements` by `source`.
**Why not reuse the old `get_ui_tree`/`ui_tree` tool-name check**: There is
no such tool; it was never callable. Removing the dead branch entirely
rather than keeping it alongside the new path avoids two divergent, only
one of which is reachable, code paths for the same concept.
### D2: `host_agent/web/app.py::_ui_tree_nodes()` reads the new field directly
**Decision**: `_ui_tree_nodes(record)` becomes a direct, unconditional read
of `record.get("ui_tree_results", [])` filtered to dicts — the same shape
as the existing `_ocr_results(record)`. The tool-name filter is removed.
**Why**: `task_detail.html` already renders `step.ui_tree_nodes` as a
collapsible list identical in structure to `step.ocr_results`; no template
change is needed once the data source is correct.
### D3: Overlay is rendered entirely client-side from data already in the page
**Decision**: The task-detail page already embeds each step's OCR/UI-tree
results and screenshot as inline data (base64 image `src`, and the
per-step context dict rendered into the template). Add a `<template>` or
inline JSON blob per step with `{elements: [...], scene: {width, height}}`
already available from `_timeline_step_context()`, and a toggle checkbox
(one global control, applying to all steps) that adds/removes a CSS class
showing/hiding an absolutely-positioned overlay `<div>` layer per
bounding box. Box position/size are computed in JS as
`left = box.x / scene.width * renderedImageWidth`, and analogously for
`top`/`width`/`height`, recomputed on image load and on window resize.
**Why client-side rather than server-rendered overlay images**: The boxes
and coordinates are already present in the page; server-side image
compositing would require a new image-processing dependency and a new
per-step render step for a purely presentational feature. Client-side
positioning also lets the operator toggle without a page reload.
**Why one global toggle rather than per-step**: Operators reviewing a task
want a consistent view across all steps while scanning; a per-step toggle
adds UI clutter for no clear benefit. (A follow-on can add per-source
toggles — OCR only vs. UI-tree only — if requested; out of scope here.)
### D4: Action-effect marker/path derived from `tool_call.args`, drawn on the before-screenshot
**Decision**: For a step whose `tool_call.action` is `"tap"`, render a
marker (e.g. a small pulsing ring) centered at `(args.x, args.y)` in the
same Scene-pixel-to-rendered-pixel scaling as D3. For `"swipe"`, render a
line/arrow from `(args.start_x, args.start_y)` to `(args.end_x, args.end_y)`
with a small dot animated along the path on load, indicating direction.
For any other action, no marker is rendered.
**Why the before-screenshot and not the after-screenshot**: `tap`/`swipe`
coordinates are the input to the action, computed against the `Scene` that
was captured before the action executed — i.e., the before-screenshot.
Drawing them there is literally accurate ("this is where the AI aimed");
drawing them on the after-screenshot would incorrectly suggest the
coordinates describe the resulting frame.
**Why derive from `tool_call.args` rather than a new persisted field**:
`_append_timeline()` already persists `tool_call={"action":..., "args":...}`
for every step; `args` already contains exactly the coordinates needed. No
new Timeline field is needed for this part.
### D5: Legacy/incomplete records degrade to no overlay/animation, not an error
**Decision**: If a step's record has no `scene` (legacy timeline record —
see `runtime-task-evidence`'s existing legacy-screenshot handling) or the
`tool_call.args` lacks the expected coordinate keys, the overlay toggle and
action marker simply render nothing for that step; the rest of the page
(screenshots, OCR/UI-tree lists, result) renders exactly as it does today.
**Why**: Consistent with the repo's existing degrade-gracefully pattern for
legacy Timeline records (single-screenshot records, missing OCR) — a
presentational enhancement must never make an old record fail to render.
## Risks / Trade-offs
- **Coordinate scaling drift**: if the actual screenshot's pixel dimensions
ever differ from `scene.screen.width`/`height` (should not normally
happen — both originate from the same captured screenshot), overlay
boxes and the action marker would be positioned slightly off. Accepted:
this mirrors an existing assumption already relied upon by the AI planner
itself (tool args are specified "in Scene pixel coordinates").
Rendered-size fallback re-measures on image `load`, so a slow-loading
image doesn't produce zero-sized boxes.
- **Overlay only on before-screenshot may surprise operators expecting it
on both**: accepted per D3/D4 — showing it on the after-screenshot would
be actively misleading, not just unhelpful, since element positions may
have moved once the action executed.
- **No per-source (OCR vs. UI-tree) toggle in this iteration**: operators
who only want one or the other must visually distinguish by box style
(e.g. color/label per `source`). Follow-on work can split the toggle if
requested.
## Migration Plan
1. Ship `TimelineRecord.ui_tree_results` (additive, defaults to `[]`
existing serialized records without the field remain readable).
2. Deploy Host Agent with the updated `_append_timeline()`,
`_ui_tree_nodes()`, and `task_detail.html`/static JS.
3. Rollback: revert Host Agent to prior version; no data migration to
reverse since the field is additive and unused by older code.
## Open Questions
- None blocking implementation.