Set `index-url` under `[tool.uv]` so all consumers (local devs and the
Jenkins image) resolve from `mirrors.aliyun.com/pypi/simple` by default
instead of relying on the per-stage `UV_INDEX_URL` env var in Jenkinsfile.
Local overrides remain available via `UV_INDEX_URL=... uv sync`.
Re-run `uv lock` to rewrite package sources/URLs from `pypi.org` to
the Aliyun mirror; versions, hashes, and resolution are unchanged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Register `SUPPORTED_DRIVER_TYPES["uiautomator2"]` backed by the new
`AndroidDriver`, mirroring `WDADriver` method-for-method. tap/swipe/home use
the confirmed UiAutomator2 mobile commands (`clickGesture`, `dragGesture`,
`pressKey`); swipe converts `duration_ms` to a drag speed with a zero-distance
guard. Adds 34 mocked unit tests, an env-var-gated integration test, and
corrects the outdated "Android not registered" note in the iPhone setup doc.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Multi-stage Dockerfile: stage 1 (node:20-bookworm-slim) builds cloud-console
with vite base "/console/"; stage 2 (uv) copies dist/ to /app/console-static.
Cloud API mounts the SPA at /console via SpaStaticFiles (StaticFiles subclass
that falls back to index.html for deep-link refreshes) when the new
CLOUD_CONSOLE_STATIC_DIR env is set, and 307-redirects / to /console/. Static
files bypass bearer auth (the SPA shell is public; tokens are still required
for /v1/*). Compose enables the mount by default; local dev still uses
npm run dev + CLOUD_CONSOLE_CORS_ORIGINS.
Jenkinsfile passes mirror overrides (NODE_IMAGE, NPM_REGISTRY, UV_IMAGE,
APT_MIRROR, UV_INDEX_URL) as --build-arg, defaulting to CN mirrors
(registry.jerryyan.net, registry.npmmirror.com, registry-ghcr.jerryyan.top,
mirrors.aliyun.com) so CN builds don't time out; Dockerfile ARGs default to
official upstreams so `docker build .` still works anywhere.
Backend suite: 443 passed (-m "not integration"); cloud-console typecheck
and production build succeed with the new base path.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements the cloud-console OpenSpec change: adds GET /v1/tasks (filterable,
bounded pagination, tasks:read) and GET /v1/tasks/{id}/attempts (404 on unknown
task) to the platform SDK, with matching CloudClient methods and a closed-by-
default CLOUD_CONSOLE_CORS_ORIGINS allow-list wired through CloudControlConfig.
Ships an independent Vue 3 + Vite SPA at cloud-console/ that authenticates with
an operator-supplied bearer token held in sessionStorage, renders tasks with
attempt history, device pool, host registry, and the plugin registry with a
registration form.
Backend test suite: 438 passed (-m "not integration"); cloud-console typecheck
and production build both succeed. PostgreSQL-backed repository tests and
manual end-to-end verification remain pending external infrastructure.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Resolve the previously-open design questions in the android-driver
change by researching the appium-uiautomator2-driver docs and Appium 3
release notes:
- tap -> mobile: clickGesture
- swipe -> mobile: dragGesture (duration_ms converted to speed px/s)
- home -> mobile: pressKey (KEYCODE_HOME)
- port isolation -> appium:systemPort capability
- Appium 3 breaking changes confirmed to not affect this design
Updates design.md (Decisions/Risks/Open Questions/Migration Plan) and
tasks.md (section 1 and tasks 2.1/2.4/4.1) accordingly.
Documents moving the repo from Windows/other dev machines to macOS and
driving a real iPhone end-to-end via Appium + WebDriverAgent: Xcode/Appium
toolchain setup, WDA signing (using the extra_capabilities passthrough
already supported by WDADriverConfig), minimal real-device verification,
and starting a Runtime API that explicitly registers and connects the
device (neither the Console nor REST API currently expose a
connect/disconnect endpoint). Also covers multi-device port allocation,
common failure modes, and a completion checklist. README gains a pointer
to it under a new "Operator Guides" section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the stub Planner's fixed describe_screen/[] behavior with a real
decision-maker: AIPlanner uses native tool/function calling (Anthropic or
OpenAI, pluggable via AI_PLANNER_PROVIDER) to select exactly one grounded
action per turn, with an explicit finish_task(success, reason) tool for
completion/failure instead of an ambiguous "no tool call" signal. Default
disabled (AI_PLANNER_ENABLED=false) and additive; TaskRunner falls back to
the existing stub Planner unchanged when disabled.
Amends CONSTITUTION.md's Perception Boundary with one narrow exception:
only the AI Planner may receive the current step's raw screenshot bytes
alongside Scene, for vision-grounded coordinate grounding. Also fixes a
latent gap in TaskRunner.run(): observe/plan exceptions are now caught per
iteration and turned into a failed task with a failure_reason, instead of
propagating uncaught.
openspec change: ai-planner-runtime.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Consumes the external Subscription Platform as source of truth for skill
content; reuses skills_learning domain models (extended with KnowledgeSkill)
and workflow.skill_exec resolver. HTTP/MCP deps land in api/ per
CONSTITUTION.md; synced skills use a physically separate SQLite file
(tasks/skills.sqlite3) to preserve the skill-authoring capability boundary.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Existing tests only exercised perception.scene_builder.build_scene()
directly or substituted a FakePerceptionProvider, so the real provider
port classes had no direct coverage. Both were already correct on
manual inspection; this closes the coverage gap.
openspec: perception-provider capability, archived change device-agent-runtime-foundation
- divergence_tolerance was loaded from config but diff_flow_versions
never consulted it, so structural_divergence was always plain
sequence-equality regardless of the configured tolerance. Now
actually applied per design.md D4.
- The skill-catalog-subscription isolation guard test asserted
'skills.catalog' not in sys.modules, which is vacuously true since
that module doesn't exist anywhere yet. Replaced with a real check
against store.py's actual imports.
openspec: skill-versioning capability, archived change skill-learning-runtime
- _with_known_widgets_only used to discard the entire SemanticScene
(returning None) when filtering dangling-element-id widgets left an
empty list, losing valid page/intents. Now returns a SemanticScene
with widgets=[] instead, matching design.md's 'dropped, not a hard
failure' intent.
- Removed _SDK_UNAVAILABLE_EXCEPTION_NAMES, dead code left over from an
abandoned name-based exception-matching approach; the blanket except
Exception already maps every SDK failure correctly.
openspec: semantic-scene capability, archived change semantic-scene-runtime
- driver/registry.py never actually defined register_driver_type, so
the plugin-system's 'driver plugin registered successfully' scenario
was unreachable in production (only simulated via test monkeypatch).
Added a real implementation wired into the existing driver factory
registry.
- PluginRegistry.discover() only caught PluginValidationError/
DuplicatePluginError, so a driver-kind manifest that failed wiring
(DriverRegistryUnavailableError/PluginTargetResolutionError) aborted
the whole scan, silently skipping co-located tool/skill manifests.
Now caught and skipped per manifest instead.
openspec: plugin-system capability, archived change cloud-runtime
CollaborativeTaskRunner hand-rolled its own step loop instead of
composing runtime/task.py's TaskRunner as design.md D3 requires,
silently dropping Timeline recording, WorldModel wiring,
TaskMetadataStore sync, and the on_task_succeeded/skill-synthesis
hook. Now calls TaskRunner's shared _start_world_view/
_record_step_result helpers for that bookkeeping. Also fixes
pre_observation being reused stale across steps in a multi-step plan
instead of refreshing to the prior step's post-observation.
openspec: multi-agent-collaboration capability, archived change multi-agent-runtime
tools/ is a Hexagonal inner layer that must never depend on LLM
concerns (ADR 0002), but describe_screen_semantic.py imported
semantic.enricher, which pulls in the Anthropic client by default.
Relocated the wrapper to runtime/, which is where LLM-dependent code
is allowed to live; updated the tool registry and all test imports
accordingly. No behavior change.
openspec: semantic-scene capability, archived change semantic-scene-runtime
- pooled_devices primary key changed from device_id alone to
(host_id, device_id), so two hosts reporting the same local
device_id no longer crash sync_host_devices() with an uncaught
sqlite3.IntegrityError.
- Device gained a capability_tags field so DevicePool.sync_host_devices()
can actually populate PooledDevice.capability_tags from a real host
sync instead of always falling back to an empty list.
openspec: device-pool capability, archived change cloud-runtime
- Resume no longer infinite-loops when a crash occurs between
append_step_result() and update_run() for a BranchStep: the recorded
branch target is now read from the persisted result instead of
recomputed as None.
- Resume no longer silently discards a recorded step failure that
happened right before the crash; the run correctly ends failed.
- A failed step's own next_step_id (e.g. routing to a BranchStep that
evaluates step_result_success) is now honored instead of
unconditionally forcing the run to failed, per spec.
openspec: workflow-orchestration capability, archived change workflow-orchestration-runtime
WorldModel routed all state through a single mutable _current_task_id,
so concurrent tasks sharing one instance could corrupt each other's
state. start_task() now returns a TaskWorldView handle scoped to that
task; TaskRunner.run() threads it through as a local variable instead
of reading self.world_model implicitly. Also extracts _start_world_view/
_record_step_result as reusable TaskRunner methods for composed runners.
openspec: world-model capability, archived change world-model-runtime