---
title: How to Review an AI-Generated Test Plan
description: "A reviewer's checklist for AI-generated test plans, worked on a 3.3 V LDO rail: coverage, limit sources, ranges, safe states, a dead-board run, approval."
url: https://galoislabs.ai/blog/review-ai-generated-test-plan
author: Alex Hernandez
author_url: https://galoislabs.ai/blog/authors/alex-hernandez
published: "2026-06-12"
topic: Agents
publisher: Galois Labs
---

# How to review an AI-generated test plan: a checklist worked on an LDO rail

![A mechanical drum sequencer in side elevation, pins driving four lever followers, one pin marked by an empty callout circle.](https://galoislabs.ai/blog/figures/agents-3.light.webp)

*FIG. 1 — SEQUENCE DRUM, ONE STEP FLAGGED*

To review an AI-generated test plan, check six things in order: each requirement maps to steps at its worst-case conditions; each limit cites a requirement or datasheet line; instrument ranges and waits fit the measurement; sequencing ends in a safe state on every path; a dead board fails the plan; and approval binds to the revision you read.

A machine-written plan reads the same whether a number came from a datasheet or from nowhere. That is the reviewer's problem. This guide works the checklist through one power rail. The method applies to any plan; the closing sections show [the same work done in Galois](https://galoislabs.ai/blog/review-ai-generated-test-plan#how-to-draft-review-and-run-this-plan-in-galois-with-évariste) and how it handles approval.

## The worked example: a 3.3 V LDO rail

Take a board whose 3.3 V rail comes from a TLV75533P, the fixed 3.3 V version of TI's TLV755P 500 mA LDO, in the SOT-23-5 (DBV) package, fed from a 5 V input. EN is tied to IN, as the datasheet suggests when shutdown is not needed. Every device figure below comes from the [TLV755P datasheet](https://www.ti.com/lit/ds/symlink/tlv755p.pdf) (SBVS320D, revised September 2024).

The board's designer wrote three requirements:

- **R1.** 3V3 stays within 3.3 V ±3% from 0 to 300 mA, for inputs from 4.5 V to 5.5 V.
- **R2.** 3V3 is in regulation within 2 ms of the input reaching 4.5 V.
- **R3.** A short from 3V3 to ground damages nothing, and the rail recovers when it is removed.

Here is an illustrative plan of the kind an agent drafts from the schematic and the datasheet:

| Step             | Conditions       | Drafted limit        |
| ---------------- | ---------------- | -------------------- |
| Power-up         | Wait 500 ms      | None                 |
| Vout             | 5.0 V in, 1 mA   | 3.267–3.333 V        |
| Vout             | 5.0 V in, 500 mA | 3.267–3.333 V        |
| Load regulation  | 1 to 500 mA      | Change ≤ 30 mV       |
| Line regulation  | 3.8 to 5.5 V in  | Change ≤ 2 mV        |
| Dropout          | 500 mA           | ≤ 215 mV             |
| Current limit    | Output shorted   | ≥ 500 mA             |
| Ripple           | 500 mA, scope    | ≤ 20 mV peak to peak |
| PSRR             | 100 kHz          | ≥ 46 dB              |
| Shutdown current | EN low           | ≤ 1 µA               |
| Input voltage    | Supply readback  | 4.75–5.25 V          |

It looks thorough. Most of it does not test the requirements.

## What to check in an AI-generated test plan

Ask the plan's author, human or agent, for evidence rather than reassurance:

| Check                      | Question                                                                            | Evidence                                                |
| -------------------------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------- |
| Requirement coverage       | Is every requirement tested at its worst case, and does every step trace to one?    | Requirement-to-step matrix                              |
| Limit provenance           | Where did each number come from?                                                    | Requirement ID, or datasheet row, column and conditions |
| Ranges and settling        | Can each instrument make this measurement, and how long does the physics take?      | Range and mode per step; a source for every wait        |
| Sequencing and safe states | What order is power applied and removed, and where does every exit leave the bench? | Power order; end state after pass, fail, error, abort   |
| Do-nothing score           | Which checks pass on a board that never powers up?                                  | The plan scored against a dead board; target zero       |
| Approval                   | What exactly is approved, and by whom?                                              | Named reviewer, pinned revision, rule for edits         |

Coverage comes first because it changes everything after it; approval comes last because it records the other five.

## Does the test plan cover every requirement?

Build the matrix before reading a single limit. For each requirement, list the conditions that make it hardest to meet, then find the steps that exercise them.

| Requirement | Worst-case conditions                    | Draft steps                                         | Finding                                          |
| ----------- | ---------------------------------------- | --------------------------------------------------- | ------------------------------------------------ |
| R1          | 4.5 V and 5.5 V in, at 0 and 300 mA      | Vout at 5.0 V in                                    | No corner tested; 500 mA exceeds the requirement |
| R2          | Input and output on a scope              | Wait 500 ms                                         | A fixed wait measures nothing                    |
| R3          | Short, bounded dwell, remove, re-measure | Current limit                                       | A different quantity; no recovery check          |
| None        |                                          | Regulation, dropout, PSRR, ripple, shutdown current | Trace to a requirement or cut                    |

The draft tests the part; the requirements describe the board, and the board's own questions (corners, timing, recovery) are missing. Read the steps against the schematic too: shutdown current needs EN driven low, and on this board EN is tied to IN, so that step cannot run as written. Untraced steps cost bench time and their own limit review.

For building a plan from requirements, see [power rail validation plans](https://galoislabs.ai/blog/power-rail-validation-plan); for keeping the links once it exists, see [hardware test traceability](https://galoislabs.ai/blog/hardware-test-traceability).

## Where did each test limit come from?

Every limit has one of three origins: a requirement, a datasheet, or nothing. A datasheet limit counts only if it is a minimum or maximum, not a typical value, at the conditions the datasheet states.

| Drafted limit            | Origin                                         | Finding                                                                 |
| ------------------------ | ---------------------------------------------- | ----------------------------------------------------------------------- |
| Vout ±1% at 1 mA         | Datasheet: ±1%, junction −40 °C to 85 °C, DBV  | Stated at VIN = 3.8 V; the board test should use R1's ±3%               |
| Vout ±1% at 500 mA       | The same row, outside its conditions           | Load regulation alone takes about 30 mV of a ±33 mV window              |
| Load and line regulation | Datasheet typical (0.060 V/A, 2 mV)            | No maximum exists to test against                                       |
| Dropout ≤ 215 mV         | Datasheet maximum at 500 mA, junction to 85 °C | Valid; the plan must define the dropout point, which the table does not |
| Current limit ≥ 500 mA   | Rated output current: a guess                  | Datasheet: 560 to 865 mA, at specific conditions                        |
| Ripple ≤ 20 mV           | Nothing                                        | No requirement and no matching datasheet figure                         |
| PSRR ≥ 46 dB             | Datasheet typical                              | Typical, and untraced                                                   |

The electrical characteristics table holds IOUT at 1 mA and VIN at VOUT(NOM) + 0.5 V (or 2.0 V, whichever is greater) unless a row says otherwise. Reusing the ±1% row at 500 mA ignores load regulation, 0.060 V/A × 0.5 A or about 30 mV typical, so a healthy part can fail.

The datasheet specifies the current limit as 560 mA minimum, 720 mA typical and 865 mA maximum with the output pulled to 90% of nominal and the input at VOUT + VDO(MAX) + 0.25 V. The draft shorts the output instead. Below 0.4 × VOUT(NOM) the regulator folds its limit back, and a dead short draws 355 mA typical (section 6.3.3). A healthy part fails the drafted limit.

Two habits keep provenance alive. Put the source in the step itself, for example "Vout at 4.5 V in, 300 mA (R1, ±3%)", so it lands in the run record. And compare every window with what eats into it. At 500 mA, 50 mΩ of wiring and contact resistance drops 25 mV, three quarters of a ±33 mV window, unless the meter senses at the board.

## Are the instrument ranges and settling times right?

**Setpoints against both sets of limits.** An instrument driver's ranges protect the instrument and know nothing about the board. The TLV755P's absolute maximum input is 6.0 V and its recommended maximum 5.5 V, so input setpoints stay at or below 5.5 V, with the supply's overvoltage protection below 6.0 V as a backstop.

**Limits and modes that measure the board, not the bench.** If the bench supply's current limit is 0.6 A, a current-limit test measures the supply. It must clear 865 mA plus ground current. The datasheet's condition holds the output at 2.97 V, which means an electronic load in constant-voltage mode; a constant-current load set above the limit drags the output into foldback. Fix the meter's range (10 V for a 3.3 V rail) rather than autoranging, and count its integration time in the step's budget, as [SCPI automation with Python](https://galoislabs.ai/blog/scpi-automation-python) shows.

**Waits that come from physics.** Each wait should name the process it covers and the source of its number. Startup time is 550 µs typical, so the 500 ms wait is about 900 times longer and still says nothing about R2, which needs a scope capture with the 2 ms limit applied to the measured delay.

Thermal settling is the wait this draft misses. At R1's hardest corner, 5.5 V in and 300 mA out, the regulator dissipates (5.5 − 3.3) V × 0.3 A = 0.66 W. The datasheet gives the DBV package 100.8 °C/W on TI's evaluation board and 231.1 °C/W on the JEDEC test board: a junction rise of about 67 °C to 153 °C. At the high end, from a 25 °C bench, the junction would pass the 165 °C thermal shutdown point. The output drifts as the junction heats, so the plan must say when the reading is taken, and the corner goes to the designer before it goes to the bench.

## Is the power sequencing safe on every exit path?

Power-up order should come with a reason for each step:

1. Set the supply's voltage, current limit and overvoltage protection with its output off.
2. Leave the electronic load's input off, then enable the supply and confirm the rail is up.
3. Enable the load. The datasheet is explicit: "For constant-current loads, disable the output load until the output rises to the nominal voltage." Foldback during startup can hold the output down.

Power-down order is a datasheet question, not a convention. The reverse-current section (7.1.4) lists three conditions: the output biased while the input is not established, the output biased above the input, and a large output capacitance with the input collapsing at little or no load current. So nothing that can source, such as a second supply or an SMU, stays on the output when the input is off. With large output capacitance, drop the input while the load still draws current, then disable the load. The 120 Ω active discharge works only while there is enough input voltage, so it cannot be counted on once the input is gone.

The short-circuit test needs a stated dwell. At 5 V in, a dead short dissipates about 5 V × 355 mA ≈ 1.8 W, and the datasheet says a sustained short cycles the part between current limit and thermal shutdown. R3 asks whether the board survives that: apply the short for a stated time, remove it, re-measure Vout against R1.

Then write the end state for every way a run can stop (pass, fail, error, operator abort, lost connection) and verify it with a measurement: supply off, load off, rail below a stated voltage. [Electronic load automation](https://galoislabs.ai/blog/electronic-load-automation) covers load modes and safe defaults.

## What would a do-nothing run score?

Score the plan against a board that never powers up: an empty fixture, a supply never enabled, a meter on the wrong channel. Any check that passes proves nothing about the hardware. The target is zero, and it runs on paper:

```python title="dead_board.py"
"""Score a test plan against a board that never powers up."""
from dataclasses import dataclass


@dataclass
class Check:
    name: str
    low: float | None   # None: no lower bound
    high: float | None  # None: no upper bound
    dead: float         # what this check reads when the rail never comes up


def passes(check: Check, value: float) -> bool:
    return (check.low is None or value >= check.low) and (
        check.high is None or value <= check.high
    )


PLAN = [
    Check("vout_1ma_v", 3.267, 3.333, dead=0.0),
    Check("vout_500ma_v", 3.267, 3.333, dead=0.0),
    Check("load_reg_delta_v", None, 0.030, dead=0.0),  # 0 V minus 0 V
    Check("line_reg_delta_v", None, 0.002, dead=0.0),  # 0 V minus 0 V
    Check("ripple_mv_pp", None, 20.0, dead=0.0),       # no output, no ripple
    Check("current_limit_ma", 500.0, None, dead=0.0),
    Check("shutdown_current_ua", None, 1.0, dead=0.0),  # nothing draws current
    Check("vin_v", 4.75, 5.25, dead=5.0),               # VOLT? returns the setpoint
]

passed = [c.name for c in PLAN if passes(c, c.dead)]
print(f"dead board passes {len(passed)} of {len(PLAN)}:")
for name in passed:
    print(f"  {name}")
```

```text
dead board passes 5 of 8:
  load_reg_delta_v
  line_reg_delta_v
  ripple_mv_pp
  shutdown_current_ua
  vin_v
```

Five of eight checks pass with nothing working. They fall into three families.

**Upper bound only.** Ripple and shutdown current have no floor, so zero passes. Give limits two sides where physics allows; where it does not, gate the check on a step that proves the rail is alive.

**Differences between readings.** Regulation subtracts two readings, and 0 V minus 0 V is within any limit. Assert each reading before the difference.

**Readback instead of measurement.** Under [SCPI-99](https://www.ivifoundation.org/downloads/SCPI/scpi-99.pdf), the query form of a command returns "the current setting associated with the command" (Volume 1, section 6.2.3), while `MEASure?` configures the instrument, takes a measurement and returns the result (Volume 2, chapter 3). `VOLT?` passes whenever the setpoint is right; `MEAS:VOLT?` reads what the output is doing.

If a known-bad unit exists, run the corrected plan against it on the bench too. A plan that cannot fail a broken board is not ready to pass a good one.

## How to approve a test plan

Approval should bind to one artifact: the exact revision (commit hash or document version), the limits with their origins, the bench configuration (which physical instruments the plan's names resolve to), and the safe-state behavior. Record the answers to the six checks with it. Any later edit, including one the agent makes, needs a new approval. If the agent also drafted an instrument driver, review it as code; [declarative instrument drivers](https://galoislabs.ai/blog/declarative-instrument-drivers) explains why a constrained format keeps that review small.

## How to draft, review and run this plan in Galois with Évariste

Évariste, the agent in the Galois platform, can draft, run and report on this plan; the review stays yours. Open it from the app sidebar (Ctrl+Shift+E) beside the project and extend the starter prompt "Create a power rail validation sequence" with this board's requirements:

> Create a power rail validation sequence for the 3V3 rail (TLV75533P, 5 V in, EN tied to IN) using psu, dmm, eload and scope. R1: 3.3 V ±3% at 0 and 300 mA, at 4.5 V and 5.5 V in. R2: in regulation within 2 ms of the input reaching 4.5 V. R3: a short from 3V3 to ground damages nothing, and the rail recovers when it is removed. Keep the input at or below 5.5 V with overvoltage protection below 6.0 V, the load off until the rail is up, and the supply off before the load. Put the requirement ID and limit in every step name.

Évariste lists the instruments connected to the team's galois-edge daemons and reads their profile commands. For a load outside the library, it generates a profile from the programming manual you upload and, after your review, deploys it to the bench and binds it. The sequence lands as a draft that cannot run until you approve it.

Review it like any other plan. The reviewed version below has setup steps with outputs off, a gate on the rail, a two-sided `numeric_limit` step per R1 corner (3.201 to 3.399 V), a scope capture for R2 and a timed short with recovery for R3. Build the coverage matrix from the requirement IDs in the step names and trace each limit to its source. Ask Évariste which SCPI command each step sends, so a `VOLT?` readback cannot pass for `MEAS:VOLT?`, and look for steps with one bound. Check R3's dwell, when each R1 reading is taken, and where every branch leaves the bench. Ask for fixes in conversation or edit in the sequence builder; every change is a new version with a diff. Approve the version you read.

Run the approved sequence on an empty fixture, and a known-bad unit if you have one; both should fail. Then run it on the board through galois-edge, watching the rail live in Monitor. Dangerous commands sent from the conversation wait for your confirmation.

Every step records its value, limits, verdict, and raw command and response. Ask Évariste which steps failed or passed close to a limit, such as the 5.5 V, 300 mA corner as the junction heats, or how the board's run compares with the empty-fixture run: a step that passed both proves nothing. Answers cite their runs and notes. "Generate a test report from the last run" produces a PDF or HTML report from a LaTeX template; add the six answers in the report editor and share it to Slack.

You keep the requirements, limits, six checks and approval, plus the wiring, sense leads, short fixture and bench safety. The driver classes, sequencing loop, error handling, command logging and report script are no longer yours to write or maintain.

| Step                 | Code path (this guide)                   | Galois with Évariste                                           |
| -------------------- | ---------------------------------------- | -------------------------------------------------------------- |
| Draft                | An agent's plan table                    | R1 to R3 in plain English; Évariste drafts the sequence        |
| Drivers              | Instrument classes you maintain          | Library profiles, or one generated from the manual             |
| Coverage and limits  | Matrix and provenance table by hand      | The same tables, built from step names                         |
| Ranges and readbacks | Range, mode and query per step           | Range in the YAML; Évariste shows each step's SCPI             |
| Do-nothing score     | `dead_board.py`, then a known-bad unit   | One-bound steps, readbacks, empty-fixture and known-bad runs   |
| Approval             | Sign-off on a commit or document version | Versioned draft with diffs, approved by revision               |
| Run                  | Your test script                         | galois-edge on the bench, watched in Monitor                   |
| Record               | Your logging                             | Per-step value, limits, verdict, raw I/O, operator, DUT serial |
| Interpret            | Read the log                             | Évariste flags failures and near-limit passes, compares runs   |
| Report               | Your report script                       | Editable PDF or HTML report, shared to Slack                   |

## How Galois handles AI-drafted test sequences

Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.

A Galois sequence goes through five stages:

1. **Drafted.** Évariste writes the sequence as YAML, and it lands as a draft.
2. **Refused while a draft.** A draft cannot start a run.
3. **Approved.** The reviewer reads the YAML as a diff. Approval records the approver and the time and applies to that exact version.
4. **Re-approved after edits.** If the sequence changes after approval, runs are refused until it is approved again. A sub-sequence it calls must also be approved and in the same project. Production revisions can be locked.
5. **Recorded.** Each step records the command sent, the raw response, the measured value, the limits and the instrument; the run records the operator and the unit's serial number.

Steps name logical instruments such as `psu` and `dmm`, which the project's bench configuration resolves to physical ones. After the power-up steps, the reviewed plan for the worked example continues like this:

```yaml title="ldo_3v3_rail.yaml (excerpt)"
name: "3V3 rail, TLV75533P (R1 to R3)"
steps:
  # action steps first: psu at 4.5 V with limits set and output off, then output_on
  - name: "Gate: run R1 to R3 only if the rail is up"
    type: condition
    config:
      instrument_id: "dmm"
      command_name: "measure_voltage_dc"
      parameters: { range: "10", resolution: "0.0001" }
      operator: ">="
      threshold: 3.0
      on_true:
        - name: "Vout at 4.5 V in, 0 mA (R1, 3.3 V ±3%)"
          type: numeric_limit
          config:
            instrument_id: "dmm"
            command_name: "measure_voltage_dc"
            parameters: { range: "10", resolution: "0.0001" }
            low_limit: 3.201
            high_limit: 3.399
            unit: "V"
            comparison: "GELE"
        # three more R1 corners, then R2 and R3, each set up by action steps
      on_false:
        - name: "Rail absent at 4.5 V in (fails the gate)"
          type: numeric_limit
          config:
            instrument_id: "dmm"
            command_name: "measure_voltage_dc"
            parameters: { range: "10", resolution: "0.0001" }
            low_limit: 3.0
            high_limit: 3.6
            unit: "V"
            comparison: "GELE"

  # end state: psu output_off before eload input off (reverse current, datasheet 7.1.4)
  - name: "Disable supply output"
    type: action
    config:
      instrument_id: "psu"
      command_name: "output_off"

  - name: "Disable load input"
    type: action
    config:
      instrument_id: "eload"
      command_name: "input_state"
      parameters: { state: "OFF" }
```

On a dead board, the gate takes its `on_false` branch, that check fails, and nothing that needs a live rail runs.

The do-nothing check applies directly: each `numeric_limit` step carries `low_limit`, `high_limit` and a `comparison` that defaults to `GELE`, and an omitted bound is disabled. For characterization, a `measure` step records a reading with no verdict; runs count those separately from passes and label a measurement-only run as recorded with no assertions.

Beneath the sequence, the [instrument profile](https://docs.galoislabs.ai/reference/instrument-profiles/) bounds what a step can send through it: declared minimum and maximum parameter values are checked before a command reaches the instrument, and ramps must run through the daemon's sweep path. [LLM instrument safety](https://galoislabs.ai/blog/llm-instrument-safety) covers the guardrails, and the [MCP reference](https://docs.galoislabs.ai/agents/mcp-server/) covers bench access.

The full loop is in [AI test automation for hardware benches](https://galoislabs.ai/blog/ai-test-automation-hardware), and the [product overview](https://galoislabs.ai/product) shows the sequence lifecycle. To try it on your own bench, start with the [quickstart](https://docs.galoislabs.ai/getting-started/quickstart/).

## Frequently asked questions

### How do you review an AI-generated test plan?

Check six things in order. Every requirement maps to steps at its worst-case conditions. Every limit cites a requirement or a datasheet minimum or maximum at matching conditions. Instrument ranges, modes and waits fit the measurement. Power sequencing reaches a safe state on every exit path. A board that never powers up fails the plan. Approval binds to the exact revision you read.

### What is a do-nothing run in test plan review?

It is the plan scored against a board that never powers up. Every check that passes in that state proves nothing about the hardware. Look for limits with only an upper bound, differences between two readings, and queries that return an instrument's setpoint instead of a measurement. Fix them with two-sided limits, a gate that confirms the rail is alive, and MEASure queries.

### Can I use datasheet typical values as test limits?

Not as pass/fail limits. A typical value is not guaranteed. TI's TLV755P datasheet, for example, gives line regulation (2 mV) and load regulation (0.060 V/A) as typical values with no minimum or maximum. Use the board's requirement, or a datasheet minimum or maximum at the conditions the datasheet states, and record which one each limit came from.

### Should an AI agent run the test plan it wrote?

Not before an engineer approves it. In Galois, a machine-authored sequence lands as a draft and the platform refuses to run a draft. Approval records who approved it and applies to the exact version reviewed. If the sequence is edited afterward, it must be approved again before the next run.

### Can the agent fix what the review finds?

Yes. In Galois, ask Évariste, the agent in the Galois platform, for each fix in conversation, or edit the sequence in the sequence builder. Every change is a new version with a diff, and an edited sequence must be approved again before it runs. Apply the same six checks to the new version, including a do-nothing run on an empty fixture.
