7.1 KiB
task-scheduler Specification
Purpose
TBD - created by archiving change cloud-runtime. Update Purpose after archive.
Requirements
Requirement: Task submission enqueues a scheduled task
The system SHALL allow a caller to submit a task (a goal string, or a reference to a WorkflowDefinition, plus optional device constraints: driver_type, required capability tags) and SHALL enqueue it as a ScheduledTask with status queued, returning a stable task id the caller can poll.
Scenario: Successful submission
- WHEN a caller submits a task with a goal and no constraints
- THEN the scheduler creates a
ScheduledTaskwith statusqueued, assigns it a unique id, and returns that id to the caller without blocking for a device to become available
Scenario: Queue depth limit reached
- WHEN a caller submits a task while the queue already holds
config.max_queue_depthqueued tasks - THEN the scheduler rejects the submission with a clear error rather than accepting an unbounded backlog
Requirement: Assignment matches a queued task to an idle, constraint-matching device
The system SHALL assign a queued ScheduledTask to an idle PooledDevice (as reported by the device-pool capability) whose driver_type and capability tags satisfy the task's constraints, using a named, registrable AssignmentStrategy.
Scenario: Matching idle device available
- WHEN
assign()runs and at least one idlePooledDevicematches the head-of-queue task's constraints - THEN the scheduler selects one such device via the configured
AssignmentStrategy, transitions the task to statusassigned, and records the chosendevice_id/host_id
Scenario: No matching device available
- WHEN
assign()runs and no idlePooledDevicematches the head-of-queue task's constraints - THEN the task remains
queued(not failed), andassign()returns without error, ready to be retried on a later call
Scenario: Unknown assignment strategy configured
- WHEN
TaskScheduleris configured with anAssignmentStrategyname that is not registered - THEN the scheduler raises a clear configuration error at startup/first-assign rather than silently falling back to a default strategy
Requirement: Assignment strategies are pluggable by name
The system SHALL provide an AssignmentStrategy registry mapping a strategy name to an implementation, with a default fifo_match strategy (oldest-queued matching task first, first matching idle device), and SHALL allow a new strategy to be added by registering a name without modifying TaskScheduler's control flow.
Scenario: Default FIFO strategy orders by submission time
- WHEN two tasks with satisfiable, overlapping constraints are queued in order A then B, and one matching idle device exists
- THEN the default
fifo_matchstrategy assigns the device to task A, leaving task B queued
Scenario: Adding a new strategy requires no scheduler edit
- WHEN a new
AssignmentStrategyimplementation is registered under a new name - THEN
TaskSchedulercan be configured to use it by name alone, with no change toscheduler.py's assignment control flow
Requirement: Local dispatch executes an assignment via existing runners
The system SHALL provide a TaskDispatcher that, for an assignment whose device is owned by the local process's own host, executes the assigned task by composing the existing agent-runtime task-execution entry point (for a goal-based submission) or the workflow-orchestration workflow-execution entry point (for a workflow-based submission), without reimplementing planning/execution/retry logic.
Scenario: Dispatching a goal-based assignment
- WHEN
TaskDispatcher.dispatch()is called with an assignment for a goal-basedScheduledTaskwhose device is local - THEN the dispatcher constructs and runs a
Taskthrough the existing task-execution entry point, and updates theScheduledTask's status todoneorfailedbased on the resulting task's outcome
Scenario: Dispatching a workflow-based assignment
- WHEN
TaskDispatcher.dispatch()is called with an assignment referencing aWorkflowDefinitionwhose device is local - THEN the dispatcher runs the definition through the existing workflow-execution entry point and updates the
ScheduledTask's status based on the resulting workflow run's outcome
Requirement: Assignment and device reservation are atomic
The scheduler SHALL atomically bind a queued task to one eligible device, create a bounded lease attempt, and reserve that device so no other active task can be assigned to it.
Scenario: Scheduler assigns an idle device
- WHEN a queued task matches an idle pooled device with no active reservation
- THEN one transaction records the assigned task, owning host/device, incremented attempt, lease identifier, lease expiry, and device reservation
Scenario: Later scheduler iteration sees stale idle snapshot
- WHEN the host snapshot still reports a device idle but that device has an active assignment lease
- THEN the scheduler excludes the device from candidates for every other queued task
Requirement: Host claim transitions assigned work to dispatched
The scheduler repository SHALL allow only the authenticated owning host to atomically claim an unexpired assigned attempt and transition it to dispatched.
Scenario: Owning host claims once
- WHEN the owning host requests available work and an unexpired assigned attempt exists
- THEN exactly one request receives the assignment and its status becomes dispatched
Scenario: Concurrent claims race
- WHEN multiple requests concurrently attempt to claim the same assignment
- THEN at most one request succeeds and every other request receives no assignment or a conflict
Requirement: Expired attempts follow bounded retry policy
The system SHALL detect expired assigned or dispatched leases and SHALL either requeue the task with its reservation released or mark it failed when the configured attempt limit is reached.
Scenario: Lease expires with attempts remaining
- WHEN an active lease expires before a terminal result and the task has remaining attempts
- THEN the task returns to queued, the previous device reservation is released, and the expired attempt remains auditable
Scenario: Lease expires at attempt limit
- WHEN an active lease expires and the task has reached its maximum attempts
- THEN the task becomes failed with a lease-expiry reason and its device reservation is released
Requirement: Terminal transitions validate the active lease
The system SHALL accept a done or failed result only from the current active task attempt and lease and SHALL make repeated identical terminal reports idempotent.
Scenario: Active lease reports completion
- WHEN the active lease owner reports a terminal result
- THEN the task transitions once to done or failed and releases its device reservation
Scenario: Superseded lease reports completion
- WHEN a result references a lease superseded by expiry and retry
- THEN the result is rejected and cannot overwrite the current task attempt