## Context The repository is currently empty (only tooling config exists: `pyproject.toml`, `.gitignore`, editor/agent config dirs, and this `openspec/` planning tree). We are building the first version of **Apex Agent**, an AI-native iPhone automation platform, from scratch. The defining constraint from the proposal is that the LLM must never be coupled to a specific automation framework (Appium/WDA) — it only ever sees semantic tools and a unified `Scene` model. Everything described here is new; there is no legacy system to interoperate with. Primary stakeholder/consumer: an LLM (GPT/Claude/Qwen/DeepSeek) acting as the agent "brain," calling into this platform either through MCP tool calls or REST, to drive one or more physical iPhones running WebDriverAgent. ## Goals / Non-Goals **Goals:** - Ship a working end-to-end loop: Observe (screenshot+OCR+tree → Scene) → Think (LLM/Planner) → Act (tool call → Driver → WDA) → Observe again. - Keep the LLM-facing surface 100% driver-agnostic: no Appium/WDA/XCUIElement types ever appear in tool names, parameters, or the Scene model. - Make the `Driver` interface implementable by something other than WDA later (e.g. `AndroidDriver`) without touching `tools/`, `vision/`, `runtime/`, or `api/`. - Persist a per-task timeline (screenshot/ocr/tree/prompt/action/result) so a task can eventually be replayed or debugged. - Expose the capability layer as MCP tools (primary) and a thin REST API (secondary, useful for debugging/dashboards). **Non-Goals:** - No script recorder, no visual flow/orchestration editor, no "keystroke wizard" style macro engine. - No second driver implementation (Android/browser/Windows) in this change — only the interface must allow it. - No production task queue (Redis/RabbitMQ) or horizontal scaling — synchronous, single-process execution is sufficient for the MVP. - No PostgreSQL/MinIO — SQLite + local filesystem is sufficient for the MVP; swapping storage later must not require changing the `runtime/` or `tools/` layers. - No icon-detection model beyond a basic/stubbed implementation (`icon_detector.py` may return "not found" or use simple template matching); full CV-based icon recognition is future work. ## Decisions ### D1: Capability Layer sits between Agent Runtime and Driver, and is the only thing tools/ calls The `tools/` package (screenshot, tap, swipe, input_text, launch_app, ui_tree, describe_screen) is a thin wrapper that calls the currently active `Driver` instance obtained from `device/` (device manager/registry), and nothing else. `runtime/` (Planner/Executor) only ever calls functions in `tools/`, never a `Driver` directly. - **Alternative considered**: Let the Agent Runtime call `Driver` directly and skip the `tools/` layer. Rejected — this would make MCP tool schemas and Driver method signatures the same thing, so any future driver-specific quirk (e.g. Android needing a different tap gesture) leaks into the LLM-facing tool contract. ### D2: `Driver` is a stateless abstract interface; `WDADriver` is the only concrete implementation for this change `Driver` (in `core/driver.py`) defines `connect/disconnect/screenshot/tap/swipe/input/launch/terminate/tree/home/lock/unlock`. `core/wda_driver.py` implements it using the Appium Python client (talking to WebDriverAgent over HTTP). The driver holds no business/task state — only the live connection handle to one physical device. - **Alternative considered**: Call WDA's HTTP API directly instead of going through Appium's Python client. Deferred, not rejected — Appium client is faster to get working for the MVP; a follow-up change can swap to raw WDA HTTP calls behind the same `Driver` interface if Appium proves too heavy, without affecting any other layer. ### D3: Scene is the single perception artifact the LLM ever consumes `vision/scene_builder.py` fuses a screenshot, the accessibility/UI tree, and an OCR pass into one `Scene` object: `{ screen: {width, height}, elements: [{id, type, text, bounds, confidence}] }`. `tools/describe_screen.py` and `tools/find_text.py`/`find_icon()`-style helpers operate only on `Scene`, never on raw XCUIElement XML or raw OCR boxes. - **Alternative considered**: Expose UI tree (XML) and OCR results as two separate tool calls and let the LLM merge them. Rejected — pushes fusion logic (dedup, coordinate reconciliation) onto the LLM, which is unreliable and burns context; fusing once in `scene_builder.py` is strictly cheaper and more consistent. ### D4: Planner produces a step plan; Executor is the only thing that calls tools and retries `runtime/planner.py` takes a natural-language goal + current `Scene` and returns a small ordered list of intended steps (e.g. "find search box", "tap it", "type query"). `runtime/executor.py` turns each planned step into concrete tool calls, owns retry/backoff/wait-for-element logic, and re-invokes the Planner (or asks the LLM again) when a step fails or the Scene doesn't match expectations. - **Alternative considered**: Let the LLM emit raw tool calls one at a time with no separate Planner ("ReAct"-style, no plan object). Rejected per the proposal's explicit design goal (P-level principle from the source discussion: "AI 不容易跑飞") — an explicit plan gives the Executor something to check progress against and reduces runaway/looping behavior; the Planner step is intentionally kept lightweight for the MVP (no long-horizon planning algorithm required). ### D5: Task Memory writes one (screenshot, Scene JSON, prompt, tool call, result) record per step to local disk, indexed by task id `storage/timeline.py` + `storage/artifact_store.py` write `tasks/history//NNN.png` and `NNN.json` (Scene + prompt + tool call + result) on every Executor step. SQLite holds task/device metadata (task id, device id, status, timestamps); images and per-step JSON stay on the filesystem, not in the DB. - **Alternative considered**: Store everything (including images) in SQLite as blobs. Rejected — bloats the DB, complicates the later MinIO migration mentioned in the proposal's deferred work, and gains nothing since images are always accessed by path, never queried. ### D6: MCP Server is the primary LLM-facing surface; REST is a secondary/debug surface over the same `tools/` functions `api/mcp.py` registers each `tools/` function as an MCP tool (name, description, JSON-schema params) with no translation layer beyond argument (de)serialization. `api/rest.py` exposes the same underlying functions via FastAPI routes (`GET /devices`, `POST /devices/{id}/tap`, etc.) mainly for manual testing/dashboards, not as the primary agent interface. - **Alternative considered**: Build REST first and put MCP behind/on top of it (MCP server calls the REST API internally). Rejected — adds an unnecessary network hop and serialization layer for the primary path; both surfaces calling `tools/` directly keeps latency low and keeps `tools/` as the true single source of truth. ## Risks / Trade-offs - **[Risk]** WDA/Appium connections are flaky (device sleep, WDA process crash, USB/Wi-Fi drop) → **Mitigation**: `Driver.connect()`/`Executor` retry logic treats connection loss as a recoverable error with bounded retries and surfaces a clear `device offline` state to the Device Manager rather than silently hanging. - **[Risk]** OCR + UI tree fusion in `scene_builder.py` may produce duplicate/conflicting elements (same visual button reported by both tree and OCR) → **Mitigation**: dedupe by bounding-box overlap (IoU threshold) in the MVP, prefer the UI-tree entry when both exist since it has ground-truth bounds; document this as a known heuristic, not a solved problem. - **[Risk]** Treating the Planner as "lightweight" could let the agent loop indefinitely on a stuck screen → **Mitigation**: Executor enforces a max-step and max-retry ceiling per task, at which point the task fails loudly (recorded in the timeline) instead of looping silently. - **[Trade-off]** Synchronous, single-process execution (no queue) means only as many concurrent tasks as the process can handle in-loop, and one slow/stuck task blocks its device — acceptable for MVP given the proposal explicitly defers queueing to a later change. - **[Trade-off]** SQLite + local filesystem storage doesn't survive multi-host deployment — acceptable for MVP (single machine, small number of physical devices); the proposal already earmarks Postgres/MinIO as deferred follow-up. ## Migration Plan Not applicable — this is the first change in a new, empty repository. There is no prior system state or data to migrate, and no rollback target other than "delete the new code." ## Open Questions - Which OCR engine ships as the MVP default (PaddleOCR, as suggested in the source discussion, vs. a lighter-weight/cloud alternative) — deferred to implementation time in `tasks.md`, since it doesn't change any interface in this design. - Whether the Planner should call the same LLM the caller is using, or be allowed a separate/cheaper model for step planning — left open; MVP can start by having the caller's LLM do both planning and per-step reasoning through the MCP tools, with `runtime/planner.py` as a thin pass-through it can later replace. - Exact MCP SDK/library choice for `api/mcp.py` — to be resolved during Sprint 4 implementation, doesn't affect the capability contracts defined here.