Implements all 19 tasks of the cloud-planner-proxy OpenSpec change:
- Cloud API: cloud.planner_config (CloudPlannerConfig, load/build helpers)
reusing runtime.tool_calling_client provider clients (no new dependency
needed -- device-cloud-platform already depends on device-agent-runtime).
- Cloud API: new host-scoped POST /internal/v1/hosts/{host_id}/planner/decide
internal endpoint, reusing existing bearer auth; logs only metadata
(host id, tool name, latency, error class), never prompt/screenshot
content.
- Host Agent: new AI_PLANNER_TRANSPORT config (direct default | cloud) and
host_agent/cloud_planner_client.py::CloudProxyToolCallingClient, a
synchronous ToolCallingClient implementation (structural, not importing
runtime) that calls the new endpoint via its own httpx.Client -- avoids
bridging the async HostAgentClient across the worker-thread boundary
that AIPlanner.plan() runs in (asyncio.to_thread in lease.py).
- Host Agent wiring: create_execution_factories()/_host_agent_planner()
select the cloud-proxy client only when AI_PLANNER_TRANSPORT=cloud;
direct/unset transport is unchanged (still the default).
- Tests: 22 new tests across Cloud API config, the new endpoint, the new
client, and transport-selection wiring; full non-integration suite
(492 tests) passes with no regressions.
- Docs: docs/CLOUD_DEPLOYMENT.md documents the cloud transport, its
trade-offs, and the credential split between Host Agent and Cloud API.
proposal.md/design.md were corrected during implementation to reflect two
findings: no new anthropic/openai dependency is actually needed, and
CloudProxyToolCallingClient uses its own sync httpx.Client rather than a
new HostAgentClient method, per the thread-boundary reasoning above.
4.1 KiB
Why
Today every Host Agent must hold its own LLM provider credentials
(ANTHROPIC_API_KEY/OPENAI_API_KEY) and provider/model configuration
locally, because runtime/ai_planner.py's AIPlanner builds an
Anthropic/OpenAI SDK client directly inside the Host Agent process
(runtime/tool_calling_client.py). With the Host Agent now defaulting AI
planning to on, this means every edge deployment must independently
provision, rotate, and secure a provider secret. Centralizing provider
configuration and credentials in the Cloud Control Plane -- which already
authenticates every Host Agent for heartbeat/claim/lease/result traffic --
removes per-edge secret sprawl and gives operators one place to change
provider/model or rotate a key without touching any Host.
What Changes
- Add a new Cloud Control Plane internal endpoint that accepts an AI Planner tool-calling decision request (system prompt, user prompt, optional screenshot, tool specs, timeout) from an authenticated Host Agent, calls the configured LLM provider using cloud-held credentials, and returns the resulting single tool-call decision.
- Cloud API owns
AI_PLANNER_PROVIDER/AI_PLANNER_MODEL/provider API keys as its own configuration; these are no longer required on the Host Agent when the new proxy transport is used. - Host Agent gains a new opt-in transport setting (proxy vs. direct-to-provider)
and a new
ToolCallingClientimplementation that calls the cloud endpoint instead of constructing a local Anthropic/OpenAI SDK client. The existing direct-to-provider transport remains fully supported and is the default, so hosts that already run with a local provider key keep working unchanged. - Reuse the existing Host-scoped bearer credential (already used for heartbeat/claim/renew/result) for the new endpoint; no new auth scope.
- Cloud API does not durably persist screenshot bytes or full prompt text from proxy requests beyond the lifetime of handling the request.
Capabilities
New Capabilities
cloud-planner-proxy: Cloud Control Plane internal endpoint and configuration that proxies AI Planner LLM tool-calling decisions on behalf of authenticated Host Agents, holding provider selection and credentials centrally instead of on each edge host.
Modified Capabilities
agent-runtime(capability defined by the not-yet-archivedai-planner-runtimechange; this delta is written against that pending spec, matching the precedent set byedge-host-self-enrollmentagainst the pendingedge-host-enrollmentspec): the "Pluggable dual-provider tool-calling abstraction" requirement is extended so the tool-calling client is selectable by transport (direct-to-provider vs. cloud-proxy) as well as by provider identity, and provider credentials become optional on the Host Agent when the cloud-proxy transport is selected.
Impact
packages/cloud-platform/cloud/apps/cloud-api: new internal API route + request/response models, new provider/credential configuration. No new package dependency:device-cloud-platformalready depends unconditionally ondevice-agent-runtime(which declaresanthropic/openai), so both SDKs are already present wherever the Cloud API runs -- confirmed byuv run python -c "import anthropic, openai"succeeding in the workspace venv without any pyproject change.apps/device-host-agent: a newToolCallingClientimplementation (host_agent/cloud_planner_client.py::CloudProxyToolCallingClient, constructed in thehost_agentpackage, notruntime, to preserve the existing hexagonal boundary that forbidsruntimefrom importingcloudorhost_agent) with its own synchronoushttpx.Client(matchingHostAgentEnrollmentClient's pattern, sinceToolCallingClient.decide()runs synchronously off the main event loop), and a new transport configuration setting (AI_PLANNER_TRANSPORT).runtime/tool_calling_client.py: no changes to theToolCallingClientProtocol itself; the new implementation satisfies it structurally from outside theruntimepackage.- Docs:
docs/CLOUD_DEPLOYMENT.mdRuntime AI Planner section gains the proxy-transport configuration path and its trade-offs.