Skip to main content

Starlight Part 3: Bounded Missions, Cancellation, and Safe Retry Decisions

How sequential missions hand off results, why ambiguous failures stop execution, and what cancellation can actually guarantee.

3 min read
Starlight Part 3: Bounded Missions, Cancellation, and Safe Retry Decisions
On this page

Revised for the September 5 platform refinement
These five articles now describe the general-purpose Node.js agent platform. The reviewed main-branch checkout identifies itself as 5.0.0-alpha.2; the wire protocol remains 1.0. The latest GitHub release is still the older v1.3.4. Use the reviewed checkout for these examples, rather than assuming the alpha has been published to a package registry.

The harder part of agent coordination is deciding when to stop. A timeout can mean the work never started, that it finished but the reply was lost, or that it is still running. Retrying blindly can turn a slow operation into a duplicate side effect.

A mission is a sequence of separately routed steps

Starlight routes each step independently. Later agents can inspect preceding results through intent.context.mission.results. Mission constraints apply throughout the sequence and cannot be overridden by a step. Agents and their verifiers still need to implement the domain meaning of those constraints.

text
mission definition
  -> select agent for step 1
  -> execute and verify
  -> pass prior results to step 2
  -> execute and verify
  -> final completed / failed / cancelled report

The order-report example makes the handoff tangible: a calculation result becomes input to a report-producing step. Read-back verification checks the artifact before the workflow is considered complete.

Retry only when the previous attempt permits it

The high-level platform stops on ambiguous errors and timeouts. Explicit retry or unhandled outcomes can permit another attempt. This conservative rule prevents the coordinator from interpreting missing confirmation as proof that an action had no effect.

ObservationEngineering response
Agent explicitly cannot handle the stepConsider another eligible agent
Ambiguous error or timeoutStop and inspect the external outcome
Verifier rejects the resultReport failure; do not continue as if the step succeeded
Operator starts a recovery missionUse a new run ID and reconcile earlier side effects

Cancellation is cooperative

Agents receive an AbortSignal and should check it around interruptible work. Cancellation cannot forcibly undo a write, stop arbitrary host code, or roll back an external API call. A verifier and an idempotency strategy remain necessary for actions that can outlive their caller.

Capacity belongs to the execution lifecycle

Registration capacity limits how much work an agent can accept. Timed-out remote attempts retain their capacity until settlement or disconnect, and the SDK accounts for outstanding work across reconnects. Shared resources such as a device or file still need a single owner or an external lock; capacity on two separate registrations is not a shared-resource lock.

Use the CLI in CI deliberately

bash
node bin/starlight-platform.js agents --agents examples/data-report/agents.cjs
node bin/starlight-platform.js run examples/data-report/mission.json --agents examples/data-report/agents.cjs

The fixed-output example rejects an existing artifact. Use npm run demo for a fresh demonstration path, or design your own mission outputs around the CI run ID. Do not delete evidence simply to make a rerun pass.

  • Validate mission definitions before dispatch; invalid definitions throw.
  • Keep reports even when a mission fails.
  • Choose timeouts based on the operation and its cancellation behavior.
  • Reconcile external state before rerunning side-effecting work.
  • Test failed verification and cancellation in addition to the happy path.

Continue the series

Reviewed implementation and setup

September refinement changelog

Core protocol specification

Dhiraj Das

About the Author

Dhiraj Das is an Automation Consultant with over a decade of experience building systems that expose failures, reduce flakiness, and make complex workflows repeatable. He applies that discipline to AI-agent validation, LLM testing, and postmortems.

He shares small open source utilities from real automation work, including: waitless (flaky tests), sb-stealth-wrapper (bot detection), selenium-teleport (state persistence), selenium-chatbot-test (AI chatbot testing), lumos-shadowdom (Shadow DOM), and visual-guard (visual regression).

Share this article: