24 KiB
Context
Today's Agent Runtime (agent-runtime, apex-agent-mvp, code-complete but unapplied) is exactly one loop shape: runtime/task.py's TaskRunner.run(task) repeats observe→plan→act until the (currently stub) Planner.goal_reached() returns true, a step fails, or TaskRunnerConfig.max_steps (default 20) is exceeded, for one flat Task.goal string. TaskContext (runtime/context.py) holds scenes/step_results for that one run and is discarded when it ends. Three other pending, unapplied changes extend what a single step can mean but not the shape of orchestration above it: semantic-scene-runtime enriches a Scene into a SemanticScene per step; world-model-runtime derives a per-task WorldState (current_app/current_page/variables/bounded history) from that same step loop; skill-learning-runtime synthesizes a parameterized FlowTemplateSkill from a completed task's Timeline, but explicitly does not build anything that can run a skill — "a synthesized skill is inert data until some future runner resolves parameters and drives tools/." None of these three changes, nor agent-runtime itself, gives a caller a way to express "do sub-goal A, then run skill B, then wait until condition C, then do sub-goal D" as one persisted, resumable unit — that is this change's entire scope.
Two stakeholders: a workflow author (human or a future higher-level Agent/orchestrator, out of scope here) who defines a WorkflowDefinition up front as data (not interactively via a UI — no visual editor exists or is planned in this change), and the process running WorkflowRunner, which must be able to pick a WorkflowRun back up after a restart without re-running already-completed mutating steps (a tap, input_text, or skill invocation is not safely re-runnable the way a read-only describe_screen is).
Goals / Non-Goals
Goals:
- Introduce
WorkflowDefinitionas an ordered, possibly-branching list ofWorkflowSteps of four kinds: planned-goal, skill-invocation, wait-for-condition, branch. - Introduce
WorkflowRunas the persisted, checkpointed execution record of oneWorkflowDefinitionrun, survivable across process restarts. - Provide a
WorkflowRunnerthat composes (imports, never edits or subclasses) the existingagent-runtimeTaskRunner/Planner/Executorfor planned-goal steps. - Provide the first real skill-invocation execution path: resolve a
FlowTemplateSkill's parameters against step-supplied argument values and drive the resolved tool calls throughtools/— the gapskill-learning-runtimedeliberately left open. - Provide a small, closed, registrable set of wait/branch condition kinds sufficient for common multi-stage flows (text appears, a tracked variable equals a value, elapsed time, prior step succeeded), extensible the same way the Driver Registry (
device-agent-runtime-foundation) is. - Make resumability real: a
WorkflowRunner.resume(run_id)reload continues from the last checkpointed step, never re-executing a step already marked completed.
Non-Goals:
- No visual workflow editor/UI —
WorkflowDefinitionis authored as data (constructed programmatically or loaded from a file/API a future caller provides); any UI isweb-console's domain, untouched here. - No distributed/multi-device workflow execution or cross-node coordination — that is Milestone 10 (Cloud Runtime); this milestone assumes one
WorkflowRunnerprocess driving one workflow run against one device at a time. - Does not replace or remove the existing single-goal Planner/Executor loop (
agent-runtime) — aWorkflowDefinitionwith exactly one planned-goal step is a valid but uninteresting workflow; simple one-shot tasks should keep callingTaskRunnerdirectly. - No general expression/scripting language for wait/branch conditions — only a closed, registrable set of condition kinds, each a small typed evaluator function, not an embedded interpreter or arbitrary
eval. - No generic skill-kind runner —
skill-learning-runtime'sSkill.kindalso allowsknowledgeskills (non-executable reference text); this change's skill-invocation step only executesflow_templateskills. Attempting to invoke aknowledge-kind skill is a step-definition error, not a runtime capability gap to fill here. - No retrieval-driven skill selection at workflow-run time — a skill-invocation step names an explicit
skill_id; usingskill-embedding-retrieval'sretrieve_candidate_skills(goal, top_k)to pick which skill a step should invoke is an authoring-time or future-Planner concern, not built intoWorkflowRunnerhere. - No changes to
runtime/,storage/,tools/,skill-catalog-subscription, orweb-console's files or specs.
Decisions
D1: New workflow/ package, sibling to semantic//world//skills_learning/, not folded into runtime/
workflow/ (models.py, runner.py, conditions.py, skill_exec.py, store.py, config.py) is its own package. runtime/ stays focused on the single-goal Observe-Think-Act-Observe loop shape (its established job per device-agent-runtime-foundation's layering ADR); multi-stage orchestration is a distinct concern one layer above it, matching the sibling-package precedent already set by semantic-scene-runtime's D1, world-model-runtime's D1, and skill-learning-runtime's D1.
Alternative considered: extend runtime/task.py's TaskRunner in place to accept a list of goals/steps instead of one goal. Rejected — TaskRunner is a pending, unapplied, code-complete capability (agent-runtime) with no applied baseline in openspec/specs/ to safely diff a MODIFIED delta against in this session; treating it as a stable, composed-over dependency (never edited) avoids retroactively changing another change's still-pending contract, and keeps "one flat goal, one loop" and "multi-stage workflow of possibly-heterogeneous steps" as two separate, independently testable concerns.
D2: WorkflowStep as a tagged union of four concrete step dataclasses, not one flat schema with optional fields
workflow/models.py defines PlannedGoalStep{goal: str}, SkillInvocationStep{skill_id: str, args: dict}, WaitForConditionStep{condition: ConditionSpec, timeout_seconds: float, poll_interval_seconds: float}, and BranchStep{condition: ConditionSpec, on_true: str, on_false: str} (all sharing a common step_id: str and optional explicit next_step_id: str | None for non-branch kinds — default is "fall through to the next step in WorkflowDefinition.steps order"), unioned as WorkflowStep = PlannedGoalStep | SkillInvocationStep | WaitForConditionStep | BranchStep. WorkflowRunner dispatches on isinstance/a kind discriminator to one handler function per step type.
Alternative considered: one flat WorkflowStep dataclass with optional fields for every kind (goal: str | None, skill_id: str | None, condition: ConditionSpec | None, on_true/on_false: str | None all present simultaneously). Rejected — this makes invalid states representable (a step with both goal and condition populated, ambiguous about which the runner should honor) and pushes validation into every consumer; a tagged union makes constructing a step with two kinds' fields at once a type error, not a runtime validation rule to remember and enforce.
D3: WorkflowRun checkpointed to WorkflowStore after every completed step, not only at start/end
WorkflowRunner.run(definition) creates a WorkflowRun{id, definition_id, status, current_step_id, variables, step_results} row via WorkflowStore.create_run(), then after each step completes (success or failure) calls WorkflowStore.update_run() to persist the new current_step_id/status/step_results entry before advancing. WorkflowRunner.resume(run_id) loads the persisted WorkflowRun, looks up its current_step_id in the WorkflowDefinition, and continues from there.
Alternative considered: persist only the initial definition and the final result, reconstructing "how far did we get" by scanning task-memory's Timeline/storage/task_metadata.py for related task ids. Rejected — Timeline/task-metadata know nothing about workflow-step boundaries (a workflow step is not a 1:1 mapping to a Task; a wait-for-condition or branch step produces no Task at all), so re-deriving workflow progress from task-memory records would require workflow-step-boundary metadata to be smuggled into another capability's storage; an explicit, workflow-owned checkpoint after each step is simpler and self-contained.
D4: Resumability is "reload-and-continue from last checkpoint," not "replay from step 0 assuming idempotent actions"
On resume(run_id), WorkflowRunner never re-executes a step whose result is already recorded in the persisted WorkflowRun.step_results for current_step_id's predecessors; it resumes execution starting at the step recorded as the current (in-progress or not-yet-started) one.
Alternative considered: on resume, replay the entire workflow from its first step, relying on tools/ actions being idempotent (e.g., re-tapping an already-tapped button is harmless) to make re-execution safe. Rejected — many mutating actions are not idempotent (sending a chat message twice creates two messages; launching a payment flow twice is unacceptable), so assuming idempotency across arbitrary drivers/tools is unsafe; explicit step-level checkpointing avoids re-invoking already-completed mutating steps, at the cost (see Risks) of not providing a stronger exactly-once guarantee across a crash during a single step's execution.
D5: Skill-invocation steps get a first, minimal parameter-resolution + tool-dispatch path in workflow/skill_exec.py
skill_exec.run_flow_template_skill(skill: FlowTemplateSkill, args: dict) -> list[StepResult] validates args against skill.parameters (missing required parameter or unknown extra key fails the step before any tool call is issued), substitutes {param} placeholders in the skill's stored steps (tool name + args-template, as produced by skill-learning-runtime's synthesis) with the resolved values, and calls the corresponding tools/* functions directly (the same functions Executor.execute() already dispatches to for planned-goal steps), returning one StepResult per resolved step. This is deliberately the "some future runner" skill-learning-runtime's design.md left as an open non-goal — Workflow Runtime is the first concrete need for it.
Alternative considered: leave skill-invocation steps declared-but-unimplemented in this change too (a no-op stub, like Planner.plan()'s unused world kwarg in world-model-runtime), deferring actual skill execution to yet another future change. Rejected — a workflow step kind that can never execute would make this proposal's own headline example ("run a Skill instead of re-planning from scratch") impossible to demonstrate; scoping it down instead (only flow_template-kind skills, no retry/backoff beyond what a direct tools/ call already offers, no retrieval-based selection — see Non-Goals) keeps the executable surface small without leaving it entirely unbuilt.
D6: ConditionEvaluator port + registry for wait/branch conditions, mirroring the Driver Registry pattern
workflow/conditions.py defines ConditionEvaluator (one method, evaluate(spec: ConditionSpec, *, scene, world_state) -> bool) and a registry mapping a condition kind string (scene_contains_text, world_variable_equals, elapsed_seconds, step_result_success) to an evaluator, the same "string key → pluggable implementation" shape as driver/registry.py's SUPPORTED_DRIVER_TYPES (device-agent-runtime-foundation, D3) and perception/provider.py's PerceptionProvider port (D8). A WaitForConditionStep/BranchStep's ConditionSpec{kind: str, params: dict} is looked up in this registry at evaluation time; adding a new condition kind is a one-function addition to conditions.py, not a change to WorkflowRunner's control flow.
Alternative considered: hardcode condition evaluation as an if/elif chain over a condition_type string directly inside WorkflowRunner. Rejected — this is precisely the extension-point shape the project has already standardized on twice (Driver Registry, PerceptionProvider); a third ad hoc if/elif chain for the same "pluggable-by-string-key" problem would be inconsistent with established project convention for no benefit.
D7: Branching via explicit on_true/on_false step-id targets, not a general expression/scripting language
BranchStep.condition is evaluated via the same ConditionEvaluator registry as wait steps; the runner then sets current_step_id to on_true or on_false (both required, explicit step ids present in the same WorkflowDefinition.steps) rather than falling through sequentially.
Alternative considered: support an embedded expression language (e.g., a small boolean-expression evaluator over WorldState.variables) for branch conditions, allowing arbitrary author-supplied predicates. Rejected — an embedded expression evaluator large enough to be genuinely useful risks becoming a code-injection-adjacent surface in a system that drives real devices (a crafted expression reaching into more than intended), and the proposal's own Non-Goals already scope out a general DSL; a closed, registrable set of condition kinds (D6) plus explicit goto-style step-id targets is sufficient for the "wait/branch" step types this milestone commits to, and is easy to extend later without an interpreter.
D8: WorkflowRunner composes TaskRunner as a black box for planned-goal steps; never edits or subclasses it
A PlannedGoalStep is executed by constructing a core.models.Task(goal=step.goal, device_id=run.device_id), calling TaskRunner(...).run(task) (letting agent-runtime's existing Planner/Executor/retry logic run unmodified), and reading back task.status/task.failure_reason to decide the step's StepResult. WorkflowRunner never imports or touches runtime/task.py's internals beyond this public run(task) -> Task contract.
Alternative considered: give WorkflowRunner its own lower-level loop that calls Planner.plan()/Executor.execute() directly per planned-goal step, skipping TaskRunner. Rejected — this would duplicate TaskRunner's already-implemented max-steps/retry/failure-reason bookkeeping inside workflow/, doubling the surface area that has to stay behaviorally consistent with agent-runtime's own spec; composing the existing public run(task) -> Task entry point is simpler and automatically inherits any future agent-runtime improvement (e.g., a smarter Planner) with zero change to workflow/.
D9: WorkflowStore is a new, independently-owned SQLite database file, not a new table in storage/task_metadata.py's existing database
workflow/store.py's WorkflowStore follows the same connect-per-call sqlite3 pattern as storage/task_metadata.py, but opens its own file (default workflows/workflows.sqlite3) with its own schema (workflow_definitions, workflow_runs, workflow_step_results tables) rather than adding tables to tasks/tasks.sqlite3.
Alternative considered: add workflow_runs/workflow_step_results tables directly into storage/task_metadata.py's existing database and module. Rejected — storage/ is owned by the pending, unapplied task-memory capability; adding tables/schema-migration code to it would be a de facto modification of another change's owned artifact with no corresponding spec delta to justify it, contradicting this session's explicit constraint not to touch other pending changes' files. A separate, workflow/-owned store with its own schema-creation code is fully additive and requires no coordination with task-memory's eventual implementation.
Risks / Trade-offs
- [Risk] Step-level checkpointing (D3/D4) does not provide a true exactly-once guarantee: if the process crashes during a mutating step's execution (after the tool call was issued but before the checkpoint write completes), resuming will re-execute that step, potentially double-sending a message or double-tapping an action → Mitigation: documented as an accepted limitation, not solved in this change (mirrors
task-memory's own per-tool-call durability limits); a future revision could require mutating skill/tool calls to carry a caller-supplied idempotency key threaded throughtools/, but no such mechanism exists intools/today to build on, so it is left as an Open Question rather than invented here. - [Risk]
WorkflowRunner's composition ofTaskRunner(D8) depends onagent-runtime's publicrun(task) -> Taskcontract remaining stable whileapex-agent-mvpitself is still pending/unapplied; if that change's spec evolves before archiving, this composition point could silently drift → Mitigation: treat onlyTaskRunner.run(task) -> Taskand theTask.status/Task.failure_reasonfields as the relied-upon contract (notTaskRunner's internals); cover this in an integration-style test that runs a real (not mocked)TaskRunnerinstance so any breaking drift fails a test, not silently. - [Risk] Skill-invocation execution (D5) drives
tools/with parameter values resolved from step-suppliedargs, which could contain a malformed or unexpectedly-typed value that passes schema validation but produces an unintended device action (e.g., an empty string where a search query was expected) → Mitigation: parameter validation checks presence/type againstskill.parameters' declared schema before any tool call is issued (fail the step, no partial execution), but does not attempt semantic validation of values (e.g., "is this a real search term") — that remains the same trust boundaryExecutor.execute()already has for planned-goal steps' tool-call arguments. - [Trade-off] Explicit
on_true/on_falsestep-id branching (D7) instead of a general expression DSL means a workflow author must hand-wire every branch target and cannot express a condition beyond the registered kinds (D6) without a code change toconditions.py→ Acceptable, matches this milestone's stated Non-Goal (no visual editor, no general DSL); a registrable-by-string-key extension point keeps adding a new kind cheap even though it requires a code change, not a data-only change. - [Trade-off] Requiring an explicit
skill_idonSkillInvocationStep(D5's Non-Goal on retrieval-driven selection) means a workflow cannot dynamically pick "whichever skill best matches this sub-goal" at run time — onlyskill-embedding-retrieval'sretrieve_candidate_skills()at authoring time (outside this runtime) could inform whichskill_idto hardcode into the step → Acceptable for this milestone; wiring retrieval into live step dispatch is future Planner-integration work, consistent withworld-model-runtime's andskill-learning-runtime's shared precedent of exposing a capability without yet teaching a decision-maker to use it dynamically. - [Trade-off] A new, separate SQLite file (D9) means "all workflow runs" and "all tasks" are two different databases with no foreign-key enforcement between a
WorkflowRun's planned-goal step and theTask/task_metadatarow it produced (only atask_idstring reference stored inWorkflowStepResult) → Acceptable; this mirrorsskill-learning-runtime's D6 acceptance of "two stores, composed in prose, merge is a future concern" rather than forcing a schema dependency on another pending change's not-yet-applied storage.
Migration Plan
This is purely additive; no existing module is edited:
- Add the
workflow/package:models.py(WorkflowDefinition, the fourWorkflowStepvariants,WorkflowRun,WorkflowStepResult,ConditionSpec),conditions.py(ConditionEvaluatorport + registry + the four starting condition kinds),skill_exec.py(run_flow_template_skill),runner.py(WorkflowRunner.run()/.resume()and one private dispatch method per step kind),store.py(WorkflowStore, schema creation for its ownworkflows.sqlite3),config.py(default poll interval, default wait timeout, default db path). - Wire
WorkflowRunner's planned-goal step handler to construct aTaskand call the existingruntime.task.TaskRunner(...).run(task)— import only, zero edits toruntime/. - Wire
WorkflowRunner's skill-invocation step handler toskill_exec.run_flow_template_skill, importingSkill/FlowTemplateSkilldataclass shapes fromskill-learning-runtime's plannedskills_learning/package (or, if implemented first,skill-catalog-subscription's canonicalskills/models.py) by import, not redefinition; if neither is implemented yet whenworkflow/is built, define the minimal shared shape locally with a note to reconcile once one of those changes lands (same caveatskill-learning-runtime's own migration plan already carries for the same dependency direction). - Wire
WorkflowRunner's wait/branch step handlers toconditions.py's registry; optionally readTaskContext.world/WorldState.variables(fromworld-model-runtime, if that change is applied and enabled) forworld_variable_equalsconditions — degrade to "condition never satisfied until timeout" ifWorldStateis unavailable, never raise. - Add
workflow*topyproject.toml's[tool.setuptools.packages.find].includelist; no new third-party dependency. - Add new tests (
tests/test_workflow_runner.py-style): a linear multi-step workflow (planned-goal → skill-invocation → wait → planned-goal), a branch workflow taking each path, a wait step that times out, and a resume test that constructs a freshWorkflowRunner/WorkflowStorepointed at the same db file mid-run to simulate a process restart and asserts already-completed steps are not re-executed. - Run
pytest— the full existing suite stays green with zero edits to existing test files, confirming this change touches no existing behavior. - Rollback: net-new package plus a net-new SQLite file with its own schema; reverting is
git revertof the commit(s), with no data migration of any existing store and no external system involved beyond whatever device the composedTaskRunner/tools/calls already target.
Open Questions
- Whether mutating skill-invocation/planned-goal steps need a stronger idempotency mechanism (e.g., a caller-supplied idempotency key threaded through
tools/) to close the crash-during-step-execution gap noted in Risks, once resumability is exercised against real, non-idempotent device flows — left unsolved here, no such mechanism exists intools/to build on yet. - Whether
WorkflowRun.variablesshould automatically bridge with a planned-goal step's underlyingTaskContext/WorldState.variables(so a value a Planner "remembers" mid-task is visible to a later branch/wait step in the same workflow) or stay a separate, workflow-only variable scope populated only by explicit step outputs — this change keeps them separate (no automatic bridging) and leaves richer variable-flow as a follow-on decision once real multi-stage workflows using both are observed. - Whether
skill-embedding-retrieval'sretrieve_candidate_skills()should eventually be wired intoSkillInvocationStepresolution (author supplies a goal instead of a fixedskill_id, runner picks the top candidate at execution time) rather than requiring an explicitskill_id— left to a future revision once retrieval-based selection has been exercised standalone. - Whether Milestone 10 (Cloud Runtime)'s distributed/multi-device execution will require redesigning
WorkflowRun's persistence schema for multi-node coordination (e.g., leader election over which node resumes a given run, or sharding runs by device) — this milestone's schema assumes single-process, single-node execution only, and that assumption should be revisited when Cloud Runtime is scoped.