Calibration status in automated test records: what to capture, when to block
By Alex Hernandez · · 14 min read


Every automated test record should name the instruments behind each measurement and say whether each one was in calibration when the run happened: manufacturer, model, serial number, firmware, calibration due date and status, captured at run time. Without that snapshot, nobody can show later that an overdue meter was not used, or find the results an out-of-tolerance meter touched.
Calibration status usually lives elsewhere: a calibration database or spreadsheet, plus a sticker on each instrument. A person reads the sticker; an unattended sequence reads nothing. This guide covers what ISO/IEC 17025 asks, which fields to capture, a Python pre-run check that blocks expired instruments, and how to find affected results when a calibration comes back out of tolerance.
Why does calibration status belong in the test record?
A calibration register answers "is this meter in calibration today?" A test record has to answer a different question, often months later: "was the meter that took this reading in calibration when it took it?" That question needs a join between two systems, and the join needs a key that the test record carries: the instrument's serial number, or an asset ID that maps to one.
Three moments make the case.
Before the run. An instrument past its due date, or marked out of service, should not take a measurement anyone relies on. People check stickers; scripts do not. If the check is not in the automation, it does not happen on the nights and weekends when long runs execute unattended.
At review. A reviewer wants to see, next to each result, which instrument produced it and that it was fit for use. That works only if the run record names the unit, not just "DMM 2". Bench slots get swapped: the meter in slot 2 this week may be the loaner that stood in while the original was at the calibration lab.
After an out-of-tolerance finding. When a meter comes back from calibration with as-found readings outside its specification, every result it produced since its last good calibration is in question. With the serial number on each step, finding those results is a query. Without it, the search runs through log files, lab notebooks and memory.
Registers also change: due dates get extended, statuses get corrected. A record that holds only a serial number and re-reads the register sees today's register. Capturing status as of run time preserves what the system believed, and allowed, when the measurement was made.
What does ISO/IEC 17025 require for calibrated equipment?
ISO/IEC 17025:2017, General requirements for the competence of testing and calibration laboratories, is the standard testing and calibration labs are accredited against. Most product teams are not accredited labs, but it is a clear written model of a defensible equipment record. The clauses that bear on calibration status:
| Clause | What it requires | What it means on an automated bench |
|---|---|---|
| 6.4.6 | Calibrate measuring equipment when its accuracy or uncertainty affects the validity of results, or traceability requires it | The DMM taking the reading needs calibration; a relay board switching it may not |
| 6.4.7 | Establish a calibration program, reviewed and adjusted to maintain confidence in calibration status | Intervals and due dates come from your program |
| 6.4.8 | Label, code or otherwise identify equipment so its user can readily see its calibration status | On an unattended run, software is the user, so it needs the status in machine-readable form |
| 6.4.9 | Take defective or out-of-specification equipment out of service and examine the effect through the nonconforming-work procedure | An out-of-service status must stop new runs and trigger a look backward |
| 6.4.13 | Keep equipment records: identity including firmware version; manufacturer, type and serial number; calibration dates, results and next due date or interval | These are the fields worth snapshotting |
| 7.5.1 | Technical records identify factors affecting the result and allow repeating the activity; observations are recorded when made | Capture identity and status at run time, not afterward |
| 7.10.1 | A nonconforming-work procedure, including an impact analysis on previous results | The out-of-tolerance query below starts that analysis |
| 7.11.2 | Systems that record or report data are validated before introduction, and so are changes | The pre-run check and the record schema are part of that system |
Two details are easy to miss. First, item a) of 6.4.13 puts software and firmware version inside equipment identity, so the firmware field of *IDN? belongs in the record, not only the serial number. Second, clause 7.8.4.3 says a calibration certificate or label shall not contain a recommendation on the calibration interval unless the customer agreed to one. The due date is a decision your program makes under 6.4.7 and records in your register; do not expect a certificate to supply it.
Metrological traceability itself, the "documented unbroken chain of calibrations, each contributing to the measurement uncertainty" of clause 6.5.1, belongs to your calibration provider and your calibration program. The test record's job is narrower: point each reading at the right link in that chain. ISO 17025 automated test records covers the standard's other record clauses for automated benches.
What calibration data should each test record capture?
Capture identity on every step and calibration status once per instrument per run. Steps are where readings happen, so each step names the instrument it used. Status is checked when the run starts, so it lives in a run-level snapshot that steps reference by serial number.
| Field | Example | Where it comes from | Why it matters |
|---|---|---|---|
| Manufacturer and model | Keysight Technologies, 34461A | *IDN? fields 1 and 2 | Confirms the right instrument type |
| Serial number | MY54505555 | *IDN? field 3, checked against the register | The key for every calibration join |
| Firmware | A.03.01 | *IDN? field 4 | Clause 6.4.13 a) counts firmware as part of identity |
| Asset ID | DMM-0042 | Calibration register | Ties the reading to labels and certificates |
| Calibration date | 2026-03-12 | Calibration register | Bounds the window an out-of-tolerance finding reopens |
| Due date | 2027-03-12 | Calibration register | The value the pre-run check compares against |
| Certificate number | C-2026-0311 | Calibration register | Leads to as-found data and uncertainty |
| Status at run time | active | Calibration register | Catches out-of-service and limited-use instruments |
| Check verdict and time | ok, 2026-10-05T08:14:03-07:00 | The pre-run check | Shows the check ran, when, and what it decided |
| Register version | SHA-256 of the export | The pre-run check | Shows which copy of the register the decision used |
Parse *IDN? defensively. The SCPI standard notes that "IEEE 488.2 is purposefully vague about the content of each of the four fields in the response syntax," and that the firmware field can hold one revision code or several (SCPI-99, section 4.1.3.6). Store the raw string, normalize the serial number before matching, and treat any mismatch with the register as a reason to stop, not to guess. Devices that do not speak SCPI, such as Modbus or CAN equipment, have no *IDN? query at all; give them an asset ID in the station configuration and carry that instead.
The pre-run check in the next section writes a snapshot like this one (values illustrative):
{
"checked_at": "2026-10-05T08:14:03-07:00",
"expected_end": "2026-10-05T14:14:03-07:00",
"allowed": true,
"register": { "path": "cal_register.csv", "sha256": "4be1c0…" },
"instruments": [
{
"resource": "TCPIP0::192.168.1.50::hislip0::INSTR",
"idn": "Keysight Technologies,34461A,MY54505555,A.03.01",
"manufacturer": "Keysight Technologies",
"model": "34461A",
"serial": "MY54505555",
"firmware": "A.03.01",
"cal": {
"asset_id": "DMM-0042",
"model": "34461A",
"serial": "MY54505555",
"cal_date": "2026-03-12",
"due_date": "2027-03-12",
"status": "active",
"certificate": "C-2026-0311"
},
"verdict": "ok",
"reasons": []
}
]
}Each step then carries the serial number, or an instrument ID that maps to it, and the snapshot answers the calibration question for every step in the run. Hardware test traceability shows the per-step record this snapshot sits beside.
How do I block test runs on expired instruments?
Run a preflight check before the sequence opens its instrument sessions: identify each instrument, look it up in an export of the calibration register, and refuse to start if anything fails. NIST's Office of Weights and Measures, in its traceability guidance for the legal metrology labs it recognizes, states the requirement plainly: "A regular review of these records and a system for preventing use of past due standards is also required" (NIST OWM, G-016). A preflight gate is that system, applied at the moment of use.
import csv
import hashlib
import io
from dataclasses import asdict, dataclass, field
from datetime import date, datetime, timedelta
import pyvisa
WARN_DAYS = 14
@dataclass(frozen=True)
class CalRecord:
asset_id: str
model: str
serial: str
cal_date: date
due_date: date # treated as the last usable day; your program defines this
status: str # for example "active", "limited", "out_of_service"
certificate: str
@dataclass
class Check:
resource: str
idn: str
manufacturer: str
model: str
serial: str
firmware: str
cal: CalRecord | None = None
verdict: str = "ok" # "ok", "warn" or "block"
reasons: list[str] = field(default_factory=list)
def normalize(serial: str) -> str:
return serial.strip().upper()
def load_register(path: str) -> tuple[dict[str, CalRecord], str]:
"""Register export keyed by normalized serial, plus the SHA-256 of the bytes read."""
with open(path, "rb") as f:
raw = f.read()
rows = csv.DictReader(io.StringIO(raw.decode("utf-8-sig"), newline=""))
register = {
normalize(r["serial"]): CalRecord(
asset_id=r["asset_id"],
model=r["model"].strip(),
serial=normalize(r["serial"]),
cal_date=date.fromisoformat(r["cal_date"]),
due_date=date.fromisoformat(r["due_date"]),
status=r["status"].strip().lower(),
certificate=r["certificate"],
)
for r in rows
}
return register, hashlib.sha256(raw).hexdigest()
def identify(rm: pyvisa.ResourceManager, resource: str) -> Check:
try:
inst = rm.open_resource(resource, timeout=5_000)
try:
idn = inst.query("*IDN?").strip()
finally:
inst.close()
except pyvisa.errors.VisaIOError as exc:
# An instrument that cannot identify itself blocks the run but still lands in the snapshot.
return Check(resource, "", "", "", "", "", verdict="block", reasons=[f"no *IDN? reply: {exc}"])
parts = [p.strip() for p in idn.split(",", 3)] # split at most 3 times; extras stay in firmware
parts += [""] * (4 - len(parts))
manufacturer, model, serial, firmware = parts
return Check(resource, idn, manufacturer, model, normalize(serial), firmware)
def judge(check: Check, register: dict[str, CalRecord], start: date, end: date) -> Check:
if check.verdict == "block": # already blocked by identify()
return check
cal = register.get(check.serial)
if cal is None:
check.verdict = "block"
check.reasons.append(f"serial {check.serial!r} is not in the calibration register")
return check
check.cal = cal
if cal.model.upper() not in check.model.upper():
check.reasons.append(f"register says {cal.model}, instrument says {check.model}")
if cal.status != "active":
check.reasons.append(f"register status is {cal.status!r}")
if end > cal.due_date:
check.reasons.append(f"calibration due {cal.due_date}, run may end {end}")
if check.reasons:
check.verdict = "block"
elif (cal.due_date - start).days <= WARN_DAYS:
check.verdict = "warn"
check.reasons.append(f"calibration due {cal.due_date}")
return check
def preflight(resources: list[str], register_path: str, expected_hours: float) -> dict:
"""Identify and judge every instrument; return the snapshot for the run record."""
register, digest = load_register(register_path)
checked_at = datetime.now().astimezone() # due dates are local calendar dates
expected_end = checked_at + timedelta(hours=expected_hours)
rm = pyvisa.ResourceManager()
try:
checks = [
judge(identify(rm, r), register, checked_at.date(), expected_end.date())
for r in resources
]
finally:
rm.close()
return {
"checked_at": checked_at.isoformat(timespec="seconds"),
"expected_end": expected_end.isoformat(timespec="seconds"),
"allowed": all(c.verdict != "block" for c in checks),
"register": {"path": register_path, "sha256": digest},
"instruments": [asdict(c) for c in checks],
}import json
from pathlib import Path
from bench.calgate import preflight
STATION = [
"TCPIP0::192.168.1.50::hislip0::INSTR", # DMM
"TCPIP0::192.168.1.51::hislip0::INSTR", # power supply
]
snapshot = preflight(STATION, "cal_register.csv", expected_hours=6)
run_dir = Path("run-7f3a")
run_dir.mkdir(exist_ok=True)
# Write the snapshot before deciding: a refused run is evidence too.
(run_dir / "calibration.json").write_text(json.dumps(snapshot, indent=2, default=str))
if not snapshot["allowed"]:
blocked = [i for i in snapshot["instruments"] if i["verdict"] == "block"]
raise SystemExit("run blocked: " + "; ".join(
f"{i['resource']}: {', '.join(i['reasons'])}" for i in blocked))
# Open sessions and run the sequence from here.The code is short; the decisions inside it are policy, and belong in your calibration program.
- Compare against the run's end, not its start. A soak that starts on the afternoon of the due date finishes the next morning, after it. The gate takes the expected duration and checks the end date.
- Decide what the due date means. Whether an instrument is usable on its due date is your program's call. This code treats it as the last usable day.
- Unknown means blocked. A serial number missing from the register blocks the run, the same as an expired one. Loaners and replacements are the usual cause: the bench changed and the register did not.
- Warn before you block. The 14-day warning gives the lab time to schedule calibration before the gate starts refusing runs.
- No grace periods in code. The same NIST guidance says calibration intervals should be extended only with supporting data and analysis. If an interval changes, change it in the register, where the decision is reviewed, not in a constant in the gate.
- Overrides are records. Sometimes a run must go ahead, such as an engineering run whose data nobody will rely on. Make the override an explicit field with a named person and a reason, store it in the snapshot, and send it through your nonconforming-work procedure if the data reaches a report.
- Validate the gate. Under 7.11.2 the check is part of the record system. Version it, test it against a register with expired, missing and out-of-service entries, and review changes like any procedure change.
The gate opens each instrument briefly to read *IDN?, so it runs before the test opens its own sessions. The session class in SCPI instrument automation with Python also starts every session with *IDN?, so each log line ties to one serial number and firmware revision.
What happens when an instrument fails calibration?
A calibration certificate under 17025 reports "the results before and after any adjustment or repair, if available" (clause 7.8.4.1 d). When the before results, usually called as-found data, show the instrument was outside its specification, clause 6.4.9 sends the lab to its nonconforming-work procedure, and 7.10.1 c) asks for "an impact analysis on previous results." Every reading the instrument took since its last in-tolerance calibration is suspect until someone reviews it.
With the serial number on every step, the impact analysis starts with one query. Assuming step results in a SQL table:
-- Every step meter MY54505555 measured since its last in-tolerance calibration
SELECT r.run_id, r.dut_serial, r.sequence_revision,
s.step_name, s.value, s.unit, s.low_limit, s.high_limit, s.verdict, s.measured_at
FROM step_results AS s
JOIN runs AS r ON r.run_id = s.run_id
WHERE s.instrument_serial = 'MY54505555'
AND s.measured_at >= '2025-03-11' -- previous calibration, found in tolerance
AND s.measured_at < '2026-03-09' -- sent for calibration; as-found out of tolerance
ORDER BY s.measured_at;Because each row holds the measured value and its limits, not only a verdict, the review can re-judge each result. Take the as-found error for the function and range the step used, subtract it from the reading, and compare again:
def rejudge(value: float, low: float, high: float, as_found_error: float) -> str:
"""as_found_error: instrument reading minus reference value, from the certificate."""
corrected = value - as_found_error
return "pass" if low <= corrected <= high else "fails when corrected"How much margin is enough, and whether to add measurement uncertainty as a guard band, is a metrology decision for your program. Each affected result then gets a recorded decision: accept, retest the unit, or notify the customer, which clause 7.10.1 e) covers along with recalling work.
Intermediate checks under clause 6.4.10, such as measuring a reference at the start of each shift and storing it as a step, can show when the drift began and narrow the window to the runs after the last good check.
Where should calibration data live?
In the calibration system, with a snapshot in each test record. Copying the whole register into the test system creates a second register that drifts. Storing only a serial number loses the run-time view. The split that holds up:
- Calibration system: owns intervals, due dates, status, certificates, as-found and as-left data, and each asset's history. It answers clause 6.4.13.
- Test record: owns identity per step and the status snapshot per run, including the verdict, the time of the check and the register version. It answers clause 7.5.1 for each result.
- The join: serial number or asset ID, normalized the same way on both sides.
Reconcile both directions periodically. Serial numbers in test records but not in the register are unregistered instruments. Active register assets no bench has reported in months are stale entries.
When is a calibration sticker enough?
A sticker and a person who checks it hold up when one engineer runs attended tests on one bench, the instruments rarely move, and no customer or assessor will ask for the record. The label is what clause 6.4.8 describes, and an automated gate adds little while someone reads the sticker before every run.
The sticker stops being enough when runs are unattended, when instruments move between benches, or when someone outside the team will review the results. That usually coincides with the move from bring-up to verification described in EVT, DVT and PVT testing.
How does Galois record instrument identity?
Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.
Galois records instrument identity at two levels. The galois-edge daemon queries each SCPI instrument with *IDN? during discovery, which runs at startup and on a configurable rescan interval, and its instrument record holds the raw *IDN? reply with manufacturer, model, serial number and firmware fields (daemon API). Each step of a Galois run records the instrument it used, the raw command and response, the measured value and its limits, and the run carries the operator, DUT serial and timestamps (product). A draft sequence cannot run until it is approved, and an approved sequence that is edited must be approved again.
Calibration dates, due dates and status belong in your calibration system. Join them to each run on serial number, read either from the daemon's instrument record or from an *IDN? snapshot your preflight takes at run start, as described above. Galois runs in Galois Cloud, a dedicated single-tenant cloud, or fully on-prem and air-gapped (deployment options), and the security page describes its audit log. The next section runs this guide's checks with Évariste, the agent in the Galois platform.
How to run these checks in Galois with Évariste
Check the bench before each run. Open Évariste from the app sidebar (Ctrl+Shift+E) beside the bench's project and ask:
List connected instruments on this bench and query
*IDN?on each. Then add an identity step per instrument at the start of the rail sequence that fails unless its reply equals the one you read.
Évariste returns each reply with its serial and firmware. Check them against a current export from your calibration system, which stays the source of truth for status and due dates; the decision to start is yours.
Review and approve. The new draft version adds one string_value step per instrument that calls the profile's identify command (*IDN?) and compares the whole reply with the string Évariste read, Keysight Technologies,34461A,MY54505555,A.03.01 for the meter. A loaner in the slot fails it, and so does a firmware update, part of identity under clause 6.4.13 a), until you approve a version with the new reply. Check the serials in those strings against the register, confirm the identity steps come first, and read the diff. Nothing runs until you approve it, and production versions can be locked (reviewing a generated test plan). Under clause 7.11.2 the identity step is part of the record system, so validate it: run it once with a different meter in the slot and keep that failed run.
Run and record. Start the run from Évariste or the app; galois-edge executes it while Monitor shows the channels live. Each step stores its raw command and response, so every run keeps the full *IDN? reply, and a different meter answering fails its identity step and marks the run failed. Then ask Évariste to "Generate a test report from the last run" and, in the report editor, add the register rows you checked and the time of the check: the run-time snapshot that calibration.json holds in the code path.
Find affected results. After an out-of-tolerance finding, ask: "List every step meter MY54505555 measured between 2025-03-11 and 2026-03-09, with run, value, limits, function and range." Évariste answers from the project's runs, citing each. Check that the list is complete. Apply the as-found corrections from certificate C-2026-0311 and decide each result yourself, or keep rejudge.py for that.
You no longer maintain run_rail_test.py, impact.sql or a report script; the due-date gate in calgate.py and the re-judgment in rejudge.py stay with your calibration process. What stays yours: the register and its policy, including what a due date means and who may override it; checking each instrument's status in the calibration system before a run; review and approval; wiring and confirming dangerous commands; and every decision on affected results under clause 7.10.1. Agent-driven test automation covers the wider loop.
| Step | Code path (this guide) | Galois with Évariste |
|---|---|---|
| Identify instruments | identify() parses each *IDN? reply | Évariste queries *IDN? on each instrument |
| Check the register | load_register() and judge() before every run | Your calibration system, compared with the replies Évariste read |
| Catch a swapped instrument | An unknown serial blocks the run | The identity step fails, and so does the run |
| Record identity | Serial number on every step | Full *IDN? reply in every run |
| Record status at run time | calibration.json, written before deciding | Register rows you checked, added to the run's report |
| Change control | Version, test and review the gate | Draft, approval, versioned diffs, production lock |
| Run | run_rail_test.py | Started from Évariste, executed by galois-edge |
| Find affected results | impact.sql and rejudge.py | Évariste lists the affected steps with citations; you re-judge |
| Report | Your report tooling | "Generate a test report from the last run" |
For the full chain from requirement to sign-off, read hardware test traceability. For the document a reviewer reads, see the DVT test report template. Whatever tools you use, the check is the same: pick one reading from last quarter and ask whether the record alone says which instrument took it and whether that instrument was in calibration on the day.
Frequently asked questions
- What is calibration status in a test record?
- It is a snapshot, taken when the run starts, of each instrument's calibration state as your calibration register reports it: manufacturer, model, serial number and firmware, calibration date, due date, certificate number and status, plus when the check ran and what it decided. Each step names the instrument it used by serial number, so every result ties to a calibration. The register itself, with intervals, due dates and certificates, stays in your calibration system.
- Does ISO/IEC 17025 require calibration status in test reports?
- Not as a listed report field: clause 7.8.2.1, which lists what a report must contain, does not name equipment. The requirement sits in equipment records (6.4.13: identity, serial number, calibration dates, due date) and technical records (7.5.1: enough information to identify factors affecting the result and to repeat the activity). Which instrument took a reading, and its calibration state at the time, is one of those factors, so the run record is the natural place for it.
- Can I use an instrument on its calibration due date?
- Your calibration program decides, and should write down, whether the due date is the last usable day or the first unusable one. Encode that rule once in the pre-run check. For long runs, compare the due date with the run's expected end rather than its start, so a soak that crosses midnight cannot finish on an expired instrument.
- What should happen when an instrument is found out of tolerance?
- ISO/IEC 17025 clause 6.4.9 takes it out of service and sends the lab to its nonconforming-work procedure, which under 7.10.1 includes an impact analysis on previous results. In practice: query every step that used that serial number since its last in-tolerance calibration, re-judge the stored values with the as-found error from the certificate, and record a decision for each affected result.
- Can I check instrument identity before every run without writing Python?
- Yes. In Galois, ask Évariste to list the instruments on the bench, query *IDN? on each, and add an identity step per instrument at the start of the sequence that fails unless the reply equals the one it read. The change lands as a draft version that runs only after an engineer approves it, so check the serials against your calibration register first. The run executes on the bench through galois-edge; a swapped instrument or a firmware change fails its identity step and the run, and each step keeps its raw command and response. Status and due dates stay in your calibration system: add the register rows you checked to the report Évariste generates.
Bring Galois to your bench.
The daemon is Apache-2.0, free forever. Enterprise runs in your cloud or on-prem.