---
title: "Test Sequencer vs AI Agent: Where Each Fits"
description: Test sequencers like TestStand and OpenTAP run the steps an engineer wrote. Test agents write, adapt and interpret them under approval. Where each fits.
url: https://galoislabs.ai/blog/test-sequencer-vs-test-agent
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-07-04"
topic: Agents
publisher: Galois Labs
---

# Test sequencer vs AI agent: where TestStand and OpenTAP stop and agents begin

![A mechanical drum sequencer in side elevation, pins driving four lever followers, one pin marked by an empty callout circle.](https://galoislabs.ai/blog/figures/agents-3.light.webp)

*FIG. 1 — SEQUENCE DRUM, ONE STEP FLAGGED*

A test sequencer executes a test an engineer authored: steps, limits and flow control are fixed before the run, and every run follows them. A test agent authors the test, adapts it when the requirement or bench changes, and interprets the results, while an engineer approves each version before it runs. Execution stays deterministic either way.

So the comparison is about roles more than products. NI now ships an agent inside TestStand, and agent platforms still hand an approved sequence to an executor. The useful questions are which work moves to the agent and what stays fixed.

> **Disclosure**
>
> We build Galois, which sits on the agent side of this comparison. Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record. Claims about TestStand, OpenTAP and PathWave link to NI's and Keysight's own documentation.

## What does a test sequencer do?

A test sequencer, or test executive, runs a predefined list of steps against instruments and a unit under test. Each step does one thing (set a supply, take a reading, compare it with limits), and rules the author set decide what happens next: continue, branch, loop or stop. The same sequence produces the same steps every time, so a pass on unit 400 means what it meant on unit 4.

**NI TestStand** is the reference design. NI describes it in one line: "Drag and drop to add and reorganize test steps, and implement conditional logic to modify sequence flow." Steps call code modules through adapters for LabVIEW, C/C++, .NET and Python. TestStand generates reports in HTML, XML, ATML and ASCII text, saves results to databases, starts steps "once equipment is available," and packages sequences to "deploy sequences to your entire fleet of testers" ([NI, What is TestStand](https://www.ni.com/en/shop/electronic-test-instrumentation/application-software-for-electronic-test-and-instrumentation-category/what-is-teststand.html)).

**OpenTAP**, an open-source sequencer under MPL-2.0 ([opentap/opentap](https://github.com/opentap/opentap)), states the model plainly: "A test plan is a sequence of test steps and their associated data," stored as XML with the `.TapPlan` extension. Instruments and DUTs are resources. Each step sets a verdict (NotSet, Pass, Inconclusive, Fail, Aborted or Error), and three break conditions, Break On Error, Break On Fail and Break On Inconclusive, can be toggled independently to decide when a plan stops. Result listeners "are notified whenever a test step generates log output, or publishes results," which is where logging and databases attach ([OpenTAP user guide](https://doc.opentap.io/User%20Guide/Introduction/Readme.html)).

**Keysight PathWave Test Automation** "leverages OpenTAP open source test automation sequencing engine" and adds Keysight's graphical tools: KS8400B for development, KS8000B as the deployment runtime, and KS8500B, a suite to "create, collaborate, deploy and monitor at scale" ([Keysight](https://www.keysight.com/us/en/products/software/pathwave-test-software/pathwave-test-automation-software.html)). [PathWave Test Automation alternatives](https://galoislabs.ai/blog/pathwave-test-automation-alternatives) compares the options.

The engines themselves do not decide what to test, write the steps, or explain a failure. Break On Fail stops the plan; it does not ask why the reading was low. Everything before the first step and after the report is engineering time, and that is where agents move in.

## What is a test agent?

A test agent is software built around a language model that turns an engineering objective into test artifacts and reads the results back. Three verbs separate it from a sequencer:

- **Author.** From a requirement, a specification or a bill of materials, it drafts the sequence and any missing instrument drivers.
- **Adapt.** When a limit, a unit or an instrument changes, it redrafts the affected steps.
- **Interpret.** After a run, it reads per-step results against limits and earlier runs and explains what changed.

"Under approval" makes an agent usable on a bench: each draft is a proposal, and an engineer approves a version before it runs.

Agents reach instruments in two patterns. In the interactive pattern, the agent calls instrument tools over the Model Context Protocol (MCP) as it works, which fits bring-up and exploration. Keysight's MCP Server for Instrument Control works this way: it resolves SCPI from each instrument's definition file, and the engineer approves each command sequence before it runs ([Keysight MCP Server](https://helpfiles.keysight.com/kmsic/English/keysight_mcp_for_instrument_control/Content/overview.html)). Session state is held in memory and cleared on restart ([release notes](https://helpfiles.keysight.com/kmsic/English/keysight_mcp_for_instrument_control/Content/release-notes.html)). In the authored pattern, the agent writes a sequence and an executor runs it. Verification work, where results must compare across units and revisions, belongs there. [AI test automation for hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware) walks through the full loop.

## Test sequencer vs AI agent, side by side

The left column describes TestStand and OpenTAP from their documentation. The right column uses Évariste, the agent in the Galois platform, as the worked example of an agent under approval.

| Dimension        | Classic sequencer                                     | Agent under approval                                                                                              |
| ---------------- | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| Authoring        | Engineer builds each step                             | Agent drafts steps, limits and drivers; engineer reviews                                                          |
| Execution        | Authored steps in order, every time                   | Same: the approved version runs as written                                                                        |
| Evidence         | Verdicts, values, reports, plus what code modules log | Per-step record with the raw command and response                                                                 |
| Change control   | Source control and review, run as a team process      | Approval gates in the run path; every edit is a diff                                                              |
| On a failed step | Break or branch rules; verdict recorded               | Same at run time; then the engineer can ask the agent about the failure and get an answer that cites earlier runs |
| New instrument   | Engineer writes or wraps a driver                     | Agent matches or drafts a profile; engineer reviews                                                               |
| Engineer's time  | Writing and maintaining steps                         | Stating objectives, reviewing, approving                                                                          |

The products sit at different points on that line:

| Tool                 | Who authors                                                                              | What executes                                                  | What is kept                                             |
| -------------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------- | -------------------------------------------------------- |
| NI TestStand         | Engineer, in the sequence editor                                                         | TestStand engine                                               | Reports and database logging                             |
| TestStand with Nigel | Nigel drafts from a specification; engineer edits                                        | TestStand engine; Nigel can start, stop, pause and resume runs | TestStand's reports and logging                          |
| OpenTAP and PathWave | Engineer, as plugins and test plans                                                      | OpenTAP engine                                                 | Whatever result listeners are installed                  |
| Keysight MCP Server  | Engineer prompts a third-party AI client; the server resolves SCPI from definition files | Command sequences the engineer approves                        | Session log, cleared on restart; saved command sequences |
| Galois               | Évariste drafts sequences and drivers; engineer approves                                 | Approved versions run through galois-edge                      | Retained per-step trace and reports                      |

## Authoring: who writes the sequence?

On a classic sequencer, an engineer authors every step: in TestStand, steps in the sequence editor calling code modules; in OpenTAP, plugins assembled into a test plan. The requirement the test is meant to prove lives elsewhere, in a specification or an engineer's memory, and the sequencer cannot check that the steps match it.

An agent starts from that requirement. In Galois, Évariste writes sequences as versioned YAML using typed steps (Numeric Limit, String Value, Pass/Fail, Action, Measure, Wait, Loop, Condition and Sequence Call) and six limit comparison operators. Steps call named commands from an instrument profile, not raw SCPI strings, so Évariste works inside a vocabulary someone has already reviewed. For an instrument outside the library of 573 profiles, an engineer uploads its programming manual and Évariste drafts a typed profile, which the engineer reviews before it is deployed to the bench; [declarative instrument drivers](https://galoislabs.ai/blog/declarative-instrument-drivers) explains why profiles are data rather than code.

Spec-driven authoring now exists inside a classic sequencer too. NI says that "In TestStand 2026 Q3, Nigel can create sequences using a specifications document that you provide," and that "It can map existing code modules or create placeholder code modules that you can edit to complete your tests" ([NI, Nigel](https://www.ni.com/en/shop/software-portfolio/nigel.html)). The difference between Nigel and an agent platform is less about drafting than about where the draft runs and what record it leaves.

Drafting is not the hard part; review is, which is why [reviewing an AI-generated test plan](https://galoislabs.ai/blog/review-ai-generated-test-plan) is its own skill.

## Execution: does an agent replace the sequencer?

No, and it should not try. A test earns trust through repetition: same steps, same order, same limits. A language model choosing the next instrument command at run time can choose differently on unit 5 than on unit 4, which is acceptable at a bring-up bench and unacceptable in a verification run.

So agent platforms keep a sequencer-shaped executor. In Galois, a run loads the approved version of a sequence and executes it as written through the galois-edge daemon on the bench. Évariste can start a run, but it does not choose commands during it: its work comes before the run (authoring) and after it (interpretation). The daemon checks typed parameter ranges from each instrument profile before a command reaches the instrument, and long ramps run as daemon-resident sweeps that survive a dropped agent session ([agents overview](https://docs.galoislabs.ai/agents/)).

For interactive work, the daemon exposes each instrument as typed MCP tools. A client that connects to the daemon directly gets no per-call auth: the port is exposed on the tailnet address and on `0.0.0.0`, so network reachability is the boundary. When an engineer sends individual commands through Évariste, a command flagged as dangerous waits for the engineer's confirmation. [LLM instrument safety](https://galoislabs.ai/blog/llm-instrument-safety) and the [MCP server reference](https://docs.galoislabs.ai/agents/mcp-server/) cover the layers in more detail.

Classic sequencers are mature here. NI lists "High-speed parallel test sequence execution" among TestStand's features ([NI TestStand](https://www.ni.com/en-us/shop/product/teststand.html)), and an agent adds nothing to a cycle time that is already tuned.

## Evidence: what does each one record?

A sequencer's evidence is its result record: verdicts, measured values against limits, and a report or database rows. How much of the instrument conversation lands in it depends on what your code modules log.

Agents raise the bar, because every run now has two accounts: what the agent says happened and what the instruments did. Only the second is evidence; a chat transcript is not a test record. In Galois, each step records the measured value, its limits, the raw command sent, the raw response received and the instrument, and each run records the operator, DUT serial and timestamps ([product](https://galoislabs.ai/product)). Pass and fail come from the limits in the approved sequence.

Interpretation sits on top of that record. Évariste reads that record: an engineer can ask which steps failed or passed close to a limit, or how a run compares with earlier units, and get an answer that cites the runs ([workplace intelligence](https://galoislabs.ai/workplace-intelligence)). The analysis sits beside the verdict and does not replace it. [Hardware test traceability](https://galoislabs.ai/blog/hardware-test-traceability) covers what a complete record needs.

## Change control: who approved this version?

With a sequencer, change control is a process the team runs around the tool: sequence files go into source control, and review and release follow the team's procedure. File format affects that review. NI recommends binary `.seq` files "for the fastest load times" and XML "only if the file must be viewable in applications without access to the TestStand engine" ([NI, improving TestStand performance](https://www.ni.com/en/support/documentation/supplemental/08/improving-teststand-system-performance.html)). Git shows a binary file only as changed, so a line-by-line review needs the XML format or a tool that reads sequence files. OpenTAP's test plans are XML, so they diff in Git.

With an agent, the author is fast and fallible, so the gate belongs in the run path. Galois enforces two. A draft is refused with "sequence is a draft and must be approved before it can run," and an edit after approval is refused with "sequence was edited after it was approved and must be re-approved before it can run." Sequences keep version history and diffs, and production versions can be locked. An agent can propose a change at any hour; it cannot make the change count without an engineer.

## How do agents adapt when requirements or instruments change?

**A tolerance tightens.** The 3.3 V rail specification moves from ±5% (3.135 V to 3.465 V) to ±3% (3.201 V to 3.399 V). On a sequencer, an engineer finds every step that checks that rail and edits its limits. An agent drafts the edit across every affected step from the updated requirement; the diff makes the approval stale, and the new limits run once an engineer approves them.

**An instrument changes.** The bench's DMM is replaced by a different model. On a sequencer, the code modules that talked to the old meter are rewritten. With profiles, the daemon matches the new meter's `*IDN?` reply to a profile, or Évariste drafts one from the manual. Évariste redrafts the steps against the new commands, and an engineer reviews both.

**A unit fails.** On a sequencer, the step records Fail, the plan breaks or continues, and an engineer opens the report. With an agent, the engineer can ask about the record: which step failed, by how much, and whether the rail has drifted across earlier units, and the answer cites the runs it drew on. The engineer decides what to measure next. Investigations, where agents propose and run follow-up experiments themselves, are a direction Galois is building toward, not a shipped capability.

In Galois, the tolerance change and the failed unit both start as a request to Évariste, opened beside the project. For the tighter tolerance, the engineer asks it to set the 3.3 V rail limits to 3.201 V and 3.399 V. Évariste edits the steps that check the rail, the change is saved as a new version with a diff, and the sequence runs again only after the engineer reviews the diff and re-approves it. For the failed unit, the engineer asks which steps failed or passed close to a limit and how the run compares with earlier units, then asks Évariste to generate a test report from the run. The engineer no longer searches the sequence for each limit or writes a script to pull and compare old results. Stating the limits from the specification, checking that the diff caught every affected step, approving the new version and deciding what to measure next stay the engineer's job.

## When a sequencer alone is the better choice

A sequencer without an agent is the right tool more often than the agent framing suggests. Choose it when:

- **The sequence rarely changes.** A mature end-of-line test running thousands of units on one revision has already paid its authoring cost. Agents do little for a test nobody edits.
- **Cycle time is the constraint.** TestStand's parallel execution and equipment scheduling are built for throughput. An agent adds nothing between steps, by design.
- **Your stations, code modules and operator interfaces work.** If LabVIEW modules, operator screens and MES integrations depend on your current reports, the cost of change sits there.
- **Your validation package names the current toolchain.** Adding a model to the authoring path can mean revisiting that package, and sometimes that costs more than the authoring time saved.
- **Policy rules out language models near test data.** A sequencer needs no model. A self-hosted endpoint, such as on-prem vLLM in a [Galois deployment](https://galoislabs.ai/deployment), still needs sign-off.
- **You want an agent without a migration.** Nigel runs inside TestStand and drafts sequences from a specification. Access requires "a valid license for the NI software and an active software service agreement" ([NI, Nigel](https://www.ni.com/en/shop/software-portfolio/nigel.html)).

And when physics sets the pace (soak, settling, chamber time), neither a faster author nor a sharper analyst moves the schedule much.

## How to choose between a sequencer and an agent

Start with where the engineering hours go. If they go into rewriting steps for boards that keep changing, as in bring-up and [EVT and DVT](https://galoislabs.ai/blog/evt-dvt-pvt-testing), authoring and adapting are what an agent takes on. If they go into reading failures across many units, interpretation is the gain. If the sequence is stable and the line is tuned, keep the sequencer.

In every case the executor stays deterministic and the run record stays the evidence; what changes is who writes the steps and who reads the results. For the product-level comparisons, see [Galois vs TestStand](https://galoislabs.ai/compare/teststand) and [Galois vs Keysight PathWave](https://galoislabs.ai/compare/keysight). [TestStand alternatives](https://galoislabs.ai/blog/teststand-alternatives) covers OpenHTF, OpenTAP and pytest as executives, and the [agents docs](https://docs.galoislabs.ai/agents/) show how an MCP client connects to a bench.

## Frequently asked questions

### What is the difference between a test sequencer and an AI agent?

A test sequencer, such as NI TestStand or OpenTAP, executes a sequence an engineer authored: steps, limits and flow control are fixed, and every run follows them. A test agent drafts that sequence from an objective, revises it when requirements or instruments change, and interprets the results. Under approval, an engineer signs off each version before it runs, and execution stays deterministic.

### Can an AI agent replace TestStand?

An agent can take over authoring and analysis, but it should not replace deterministic execution. Something still has to run the approved steps in the same order, with the same limits, on every unit. That can be TestStand with NI's Nigel agent drafting sequences, or an agent platform with its own executor. Stable, high-volume production lines often need no agent at all.

### Does NI TestStand have an AI agent?

Yes. NI's agent, Nigel, shipped as an advisor with TestStand 2025 Q3. In TestStand 2026 Q3 it can create sequences from a specifications document you provide and map existing code modules or create placeholder ones, and it can start, stop, pause and resume a run. Its in-product actions ask for confirmation, and access requires an NI license with an active software service agreement.
