Skip to main content

Starlight Part 2: Reading Execution Evidence Instead of Trusting a Green Badge

A practical guide to the new report explorer, six recorded scenarios, and the difference between a completed action and a verified outcome.

3 min read
Starlight Part 2: Reading Execution Evidence Instead of Trusting a Green Badge
On this page

Revised for the September 5 platform refinement
These five articles now describe the general-purpose Node.js agent platform. The reviewed main-branch checkout identifies itself as 5.0.0-alpha.2; the wire protocol remains 1.0. The latest GitHub release is still the older v1.3.4. Use the reviewed checkout for these examples, rather than assuming the alpha has been published to a package registry.

A dashboard should help a reviewer inspect an outcome. The September Starlight website replaces the older browser Mission Control story with a read-only explorer of six actual execution reports and a narrated demonstration. It makes the evidence visible without pretending that a static website is controlling live agents.

Begin with the result, then inspect the attempt

Open a report and identify its mission, final status, selected agent, attempts, and evidence. A completed status is meaningful only in the context of the agent and its verifier. It does not certify that arbitrary business constraints were enforced.

ScenarioQuestion the evidence should answer
Verified outputDid the generated artifact contain the expected result?
Constraint failureWas disallowed work rejected rather than silently completed?
Rejected verificationDid a failed outcome check stop progression?
CancellationWas cancellation reported and outstanding work accounted for?
Authenticated remote executionDid a separate agent connect through the authenticated path?
Fresh recovery runWas a new run created and verified after the earlier failure?

Watch the demonstration with that checklist

The public demo is a 3 minute 20 second, 1080p narrated walkthrough with captions and chapter controls. Its purpose is to connect visible status to execution evidence, including failure paths. The explorer contains recorded examples; running a fresh mission locally is how you check the behavior in your own environment.

Inspect your own run

Run the data-report demo from Part 1, then use the run ID printed by the CLI. Replace the placeholder below with that ID.

bash
node bin/starlight-platform.js inspect <run-id>

Keep the final JSON report and its output artifact together. If a test result is important to your decision, preserve the command and input revision too. A report from a different checkout is useful history, but it cannot validate a later edit.

Measure savings instead of assigning them

The earlier article assigned fixed minutes of saved engineering work to each automated intervention. That was an estimate, not measured ROI. To assess value, compare comparable tasks before and after adoption: time to review, intervention frequency, repeated work, and recovery effort. Record the number of runs and the environment alongside the result.

Know what the report leaves out

  • CLI reports are final records, not durable execution checkpoints.
  • The SDK retains a bounded in-memory history, with a default of 100 runs.
  • A cancelled task may already have created a side effect.
  • Keep credentials and personal data out of mission context and evidence.
  • A new recovery run is a separate attempt; it does not resume a crashed process automatically.

The strongest report makes it easy to disagree with the agent. A reviewer should be able to find a missing check, rejected result, or incomplete action without reconstructing the entire session from a success message.

Continue the series

Reviewed implementation and setup

September refinement changelog

Core protocol specification

Dhiraj Das

About the Author

Dhiraj Das is an Automation Consultant with over a decade of experience building systems that expose failures, reduce flakiness, and make complex workflows repeatable. He applies that discipline to AI-agent validation, LLM testing, and postmortems.

He shares small open source utilities from real automation work, including: waitless (flaky tests), sb-stealth-wrapper (bot detection), selenium-teleport (state persistence), selenium-chatbot-test (AI chatbot testing), lumos-shadowdom (Shadow DOM), and visual-guard (visual regression).

Share this article: