EVT vs DVT vs PVT: what each build proves and what to test
By Alex Hernandez · · 14 min read


EVT (engineering validation test) proves the design works on prototype boards. DVT (design validation test) proves the production-intent design meets every requirement, including environmental, reliability and compliance tests. PVT (production validation test) proves the factory can build and test it at rate. Each stage should build on the previous stage's tests, not start over.
The names are industry conventions, and companies draw the lines differently; some add a prototype build before EVT or a ramp build after PVT. In chip work, PVT also means process, voltage and temperature, the corners that post-silicon validation bench automation sweeps. This guide covers each build in turn, then the part most plans skip: keeping one set of test assets from the first board to the production line.
What do EVT, DVT and PVT stand for?
Each build answers one question, on hardware built to answer it:
| Stage | Question it answers | Hardware | Built where | Exit when |
|---|---|---|---|---|
| EVT | Does the design work? | First full-design prototypes, sometimes in several variants | Lab or prototype shop | Every function works and the changes for DVT are known |
| DVT | Does the production-intent design meet every requirement? | Final layout and parts, tooled or near-tooled mechanics | Pilot line, ideally on the production process | Every requirement verified, compliance and reliability passed, design frozen |
| PVT | Can the line build and screen it at rate? | Production units | Production line with production fixtures and operators | Yield, capability and cycle-time targets met, production test released |
The same failure means different things at each stage. At EVT it is a design bug to find. At DVT it is a requirement miss, and the fix reopens verification. At PVT it is a process, part or test-station problem, unless it exposes a margin DVT missed. Each later build has more tooling, fixtures and records attached to the design, so a change costs more at each step.
What to test in EVT
EVT boards are the first builds of the full design, often hand-assembled and reworked on the bench. The goal is coverage: exercise every block and interface at least once, then measure how much margin the design has.
Typical EVT test content:
- Bring-up. Power-on with current limits, rail voltages and sequencing, clocks, resets, programming and boot. The board bring-up checklist covers the order.
- Functional coverage. Every interface, sensor, actuator and firmware feature exercised against its design intent, including paths that only matter under fault conditions.
- Characterization. Rails, regulators, oscillators and analog chains measured across input voltage, load and temperature corners, and compared against datasheet and requirement limits. The power rail validation plan works through one subsystem in depth.
- Early risk tests. Thermal hot spots under worst-case load, signal integrity on the fastest buses, and an EMC pre-compliance scan while layout changes are still cheap.
How many units. EVT makes no statistical claim, so count test tracks, not confidence levels. Plan one unit per parallel track (bring-up, firmware, characterization, thermal, pre-compliance), plus spares for rework and for the boards that get damaged. Each design variant needs its own tracks. A single failure is a bug to root-cause, not a data point.
Exit criteria. Every function has been exercised. Every blocking issue has a root cause and a fix proven on a reworked unit. Characterization data shows where margin is thin. The list of design changes for DVT is closed.
What to automate. Instrument drivers, which carry into DVT unchanged, and characterization sweeps, which become DVT verification tests once limits are attached. Keep fixtures simple, because the board will change.
What to test in DVT
DVT units should match production as closely as the schedule allows: final layout, final component choices, tooled or near-tooled mechanics, release-candidate firmware and, ideally, the production assembly process. DVT is where requirements are verified one by one, so every requirement needs a named test, a limit and a record.
Typical DVT test content:
- Full electrical verification. Every requirement measured against its limit across the specified supply, load and temperature corners, on several units rather than one.
- Environmental. Operating and storage temperature, humidity, thermal cycling, vibration, shock and drop, to the profiles your market or customer specifies. Automotive electronics, for example, are tested for electrical loads to ISO 16750-2, and airborne equipment to the environmental categories of DO-160.
- Reliability. Highly accelerated life testing (HALT) to find operating and destruct margins, and accelerated life tests sized to the reliability claim the product has to support.
- Compliance. Formal EMC, safety and radio tests, usually at an outside test lab. These need production-representative hardware. Under the FCC's Supplier's Declaration of Conformity, the declaration covers marketed units that are identical to the sample tested (47 CFR 2.906(b)), and identical means "within the variation that can be expected to arise as a result of quantity production techniques" (47 CFR 2.908). For products with digital elements placed on the EU market, the Cyber Resilience Act requires cybersecurity test reports in the technical documentation from December 11, 2027.
- Regression after every change. Each engineering change order reopens the tests it can affect.
How many units does DVT need?
Add up the tracks. Some tests are destructive, such as drop and accelerated life to failure. Compliance labs keep samples. Each environmental profile ties up units for days or weeks. The largest single track is often the reliability claim.
For a pass/fail claim with zero failures allowed, the arithmetic is short. If each unit passes with probability R, then n independent units all pass with probability Rn. To claim reliability R at confidence C, pick the smallest n with Rn ≤ 1 − C, which is n = ln(1 − C) / ln(R), rounded up. Allowing one failure supports the same claim but needs more units:
| Reliability claim | Confidence | Units, zero failures | Units, one failure allowed |
|---|---|---|---|
| 90% | 90% | 22 | 38 |
| 90% | 95% | 29 | 46 |
| 95% | 90% | 45 | 77 |
| 95% | 95% | 59 | 93 |
| 99% | 90% | 230 | 388 |
| 99% | 95% | 299 | 473 |
Every count comes from the binomial distribution: the smallest n for which the chance of seeing that few failures, if the true reliability were R, is at most 1 − C. The last two rows explain why high reliability claims are usually supported by accelerated testing and field data rather than unit count alone.
For a claim in time rather than pass/fail, the NIST/SEMATECH e-Handbook gives the zero-failure plan for an exponential life model: test time is the MTBF objective times a factor, which is 2.30 at 90 percent confidence. Its example confirms a 200-hour MTBF at 90 percent confidence with 460 failure-free hours, and the hours can be split across units, since the model treats one unit for 1,000 hours the same as two units for 500 each (NIST e-Handbook 8.3.1.1). Both methods assume the test conditions represent use. That assumption is the hard part to argue, and it is the first thing a reviewer will ask about.
Exit criteria. Every requirement traces to a passing test record on DVT units (hardware test traceability). Every failure has a root cause and a verified fix. Compliance reports are in hand. The reliability claim is demonstrated at the stated confidence. The design is frozen under change control. The DVT test report template lays out the deliverable.
What to automate. The full verification suite, as versioned sequences with limits taken from the requirements. Environmental runs are long and unattended, so they need logging that survives a weekend: every measurement, the unit serial, the chamber state and the instrument identity. Generate reports from those records, because each change order reruns the same suite.
What to test in PVT
PVT units come off the production line: production tooling, production fixtures, production operators and the production test station, at or near the planned rate. The design is assumed done. PVT tests the process and the test that screens it.
Typical PVT test content:
- The production test itself. In-circuit or flying-probe test for assembly faults, boundary scan where the parts support IEEE 1149.1, device programming, and a functional test of the parameters DVT showed to be sensitive.
- Measurement system analysis. Before trusting a station's numbers, characterize its error: repeatability on the same unit, and reproducibility across operators, fixtures and stations. The NIST/SEMATECH e-Handbook covers these as gauge R&R studies.
- Process capability. For key parameters, compare the spread of production units with the spec limits. Cpk is the distance from the process mean to the nearer spec limit, divided by three standard deviations (NIST e-Handbook 6.1.6).
- Correlation with DVT. Run a set of PVT units through the DVT bench suite and compare the results. If production units measure differently from DVT units, find out why before shipping them.
- Yield, cycle time and audits. First-pass yield at each station, test time against the line's takt time, and out-of-box audits on packed units.
How many units. Enough to run the line at rate until its real variation shows, and to estimate capability on each key parameter. The same NIST section says capability estimates are valid only with a large enough sample, "generally thought to be about 50 independent data values," and that capability studies generally need n ≥ 100. Add units for audits, reliability spot checks and the DVT correlation.
Exit criteria. Yield and cycle time meet the targets the team set. Capability on key parameters meets its target. The station's measurement error is small relative to the tolerance it judges. Correlation with DVT holds. Production test sequences are released and locked.
What to automate. Everything at the station, from serial scan to yield reporting. Operators should never type a limit or choose a test version.
How should test assets carry from EVT to PVT?
The expensive pattern is three test programs: a folder of EVT scripts, a DVT suite in another tool, and a production test written by the contract manufacturer from a spec sheet. Each rewrite loses what the last stage learned, and none of the results compare. Treat test assets as one product that matures with the hardware:
| Asset | EVT | DVT | PVT |
|---|---|---|---|
| Instrument drivers | Written for the bench instruments | Same drivers, plus chambers and environmental instruments | Same drivers where instruments match; station drivers checked against them |
| Test sequences | Bring-up and characterization scripts | Verification suite, one test per requirement | Production subset selected by risk |
| Limits | Observed distributions and datasheet values | Spec limits from requirements | Test limits inside spec, guard-banded for measurement uncertainty |
| Records | Serial, revision, instrument, raw readings | Same schema, plus chamber state and requirement IDs | Same schema, plus station, fixture and operator |
| Fixtures | Bench leads and probes | Repeatable fixtures for environmental runs | Production fixtures correlated against DVT |
Four rules keep the assets intact:
- One driver layer. Write each instrument's driver once, as a typed class or a declarative profile, and call it from every stage. When a measurement disagrees between DVT and PVT, the difference is then in the fixture or the unit, not the driver. The SCPI automation guide shows what a driver needs underneath.
- Limits are data, not code. Store each limit next to its sequence step, with the requirement it comes from. Tightening a limit from DVT to PVT then becomes a reviewed diff, not an edit buried in a script.
- One record schema from the first board. Record the unit serial, hardware revision, sequence version, instrument identity and raw instrument response at EVT, even when nobody asks for them yet. DVT and PVT results then compare against EVT without a conversion step.
- Derive the production test; do not rewrite it. Pick the production subset from the DVT suite by asking which parameters vary with assembly and parts, and which failures the earlier builds actually found. Then correlate the production fixture against the DVT bench on the same units, including known-bad units if you have them.
Guard banding ties rules 2 and 4 together. A production limit sits inside the spec limit by a margin derived from the station's measurement uncertainty, which is why the gauge R&R study comes before production limits are frozen, not after.
What to automate at each stage
| Stage | Automate first | Leave manual for now | Why |
|---|---|---|---|
| EVT | Instrument drivers, characterization sweeps, data capture | Fixtures and operator interfaces | The design changes weekly; sweeps are where manual time goes |
| DVT | Full verification suite, unattended environmental runs, reports | One-off failure analysis | Every change order reruns the suite; repeatability is the point |
| PVT | The whole station flow, per-unit records, yield and capability reporting | Nothing on the line | Operators work to takt time; manual steps become escapes |
The pattern is to automate what the next stage will repeat: EVT characterization becomes DVT verification, which supplies the production subset. Running hardware tests in CI covers keeping the suite running between builds.
How do you run EVT, DVT and PVT tests in Galois with Évariste?
Évariste, the agent in the Galois platform, can take one test from the first board to the line as a sequence that matures under review. Open it from the app sidebar (Ctrl+Shift+E) beside the project; it reaches the bench through galois-edge. AI test automation for hardware benches covers the agent in depth.
EVT: drivers and the sweep. Ask "List connected instruments" to confirm the DC supply, electronic load and DMM. For an instrument without a profile, upload its programming manual; Évariste generates one and, after you review it, deploys it to the edge and binds it. Then state the objective:
Create a power rail validation sequence for VCCINT: inputs 4.75, 5.0 and 5.25 V, loads 0.5 and 3.0 A with 2 s settling, VCCINT on the DMM at the FPGA test point, 0.95 to 1.05 V (1.0 V ±5%, consumer datasheet). Record each point; this is characterization, so keep the readings even where a limit fails.
This is the VCCINT row of the power rail validation plan walkthrough, which has the full prompt and shutdown order.
DVT: requirements and corners. Ask Évariste to edit the same sequence, not start a new one: take each limit from its requirement, put the requirement ID in the step name, repeat the six points at the chamber setpoints the requirement names, and record the chamber temperature and a board thermocouple at every point. A chamber controller without a profile gets one the same way. The hot corner of the new draft:
name: "VCCINT line and load regulation"
steps:
# setup: chamber to 70 °C and soak, input set to 4.75 V, load to constant
# current, outputs on; then the 0.5 A point, load to 3.0 A, settle 2 s,
# and a chamber temperature reading from the chamber controller
- name: "Board temperature at 70 °C, 4.75 V, 3.0 A"
type: measure
config:
instrument_id: "daq"
command_name: "measure_temperature"
unit: "degC"
# limit source: REQ-PWR-012, VCCINT 1.0 V +/-5% over the operating range
- name: "REQ-PWR-012 VCCINT at 70 °C, 4.75 V, 3.0 A"
type: numeric_limit
config:
instrument_id: "dmm"
command_name: "measure_voltage_dc"
low_limit: 0.95
high_limit: 1.05
unit: "V"
comparison: "GELE"
# same points at each setpoint; closing steps turn the load off, then the supplyReview and approval. Each draft waits for your approval before it runs. Check that every requirement has a step, each limit matches its requirement, the corners cover the operating range and shutdown runs in reverse. How to review an AI-generated test plan has the checklist. Every edit, in conversation or the sequence builder, is a new version with a diff to review, as rule 2 asks. A command a profile marks as dangerous, sent on its own from the conversation, waits for your confirmation before Évariste sends it.
Runs, results and report. Start the run; galois-edge executes it while Monitor shows channels live. Every step records its measured value, limits, pass or fail, raw command and response, instrument, operator, DUT serial and timestamps. Keep the build stage and hardware revision in a project note Évariste can cite. Ask which points failed or passed closest to a limit, and how DVT compares with EVT; check its answers against the recorded steps. Then ask it to "Generate a test report from the last run", edit it in the report editor and share it to Slack.
PVT: the production subset. The failed and near-limit steps across the EVT and DVT runs, plus the parameters that vary with assembly and parts, are rule 4's candidates; the choice stays yours. Ask for a production sequence with the steps you pick and limits guard-banded inside spec, such as 0.96 to 1.04 V for a 10 mV band from the gauge R&R study. Run the same PVT units on the DVT bench and at the station, have Évariste compare the runs, and production-lock the released version once correlation holds.
You no longer write or maintain driver classes, sweep loops, limit files, retry and error handling, logging, CSV wrangling, plotting or a report script. Objectives and limits from requirements and datasheets, review and approval of every version, the production subset and guard band, fixtures, wiring, and chamber and bench safety stay with you.
| Step | Code path (this guide) | Galois with Évariste |
|---|---|---|
| Drivers | Typed class or declarative profile | Library profile, or generated from the manual and reviewed |
| EVT sweep | Characterization scripts | Drafted from a plain-English objective |
| DVT limits | Limits as data, with requirement IDs | Requirement ID and limit in each step |
| Change orders | Reviewed diff | New version with a diff for you to review |
| Environmental runs | Logging that survives a weekend | Run through galois-edge; Monitor live |
| Records | One schema from the first board | Per-step value, limits, raw response, DUT serial |
| Interpret | Your analysis scripts | Failed and near-limit steps; run comparison |
| Report | Generated from the records | Generated report; report editor |
| PVT subset | Derived from DVT, guard-banded | Your chosen steps, production-locked |
| Correlation | Same units on DVT bench and station | Évariste compares the runs |
Where Galois fits in an EVT, DVT and PVT plan
Galois is agent-driven test engineering for hardware teams: agents generate tests and instrument drivers, run them on real benches through the open-source galois-edge daemon, and turn the results into reports and a shared engineering record.
Each part of the asset model maps to something in the product, as the Évariste walkthrough shows:
- Starting the EVT plan. From schematic to test plan covers drafting the first plan from the parts on the board. Galois design plugins for KiCad and Altium are in build.
- Drivers written once. The instrument library lists 573 instrument profiles, and an instrument without one gets a profile drafted from its manual for review. The same profile serves the EVT bench, the DVT chamber rack and the PVT station through the galois-edge daemon.
- Sequences that mature under review. Sequences have version history, diffs and production locking; a sequence Évariste writes or edits runs only after an engineer approves it.
- One record. Every run keeps the per-step record described above. Reports are drafted from those records (product).
Teams with an existing TestStand line can read the TestStand comparison for how sequences map. To try the daemon on your own bench, start with the quickstart.
Frequently asked questions
- What is the difference between EVT, DVT and PVT?
- EVT (engineering validation test) checks that the design works on early prototypes. DVT (design validation test) checks that the production-intent design meets every requirement, including environmental, reliability and compliance tests. PVT (production validation test) checks that the production line, its fixtures and its production test can build and screen the product at rate.
- How many units do you need for DVT?
- Add up the test tracks. A pass/fail reliability claim often sets the largest one: if each unit passes with probability R, n units all pass with probability R to the power n, so showing 90 percent reliability at 90 percent confidence takes 22 units with zero failures, and 95 percent at 95 percent confidence takes 59. Destructive tests, compliance lab samples and spares come on top.
- Can I run EVT, DVT and PVT tests in Galois without writing Python?
- Yes. Ask Évariste, the agent in the Galois platform, in plain English for the EVT sweep on the DC supply, electronic load and DMM; it drafts a sequence with limits, then edits that same sequence for DVT with requirement IDs and chamber corners, and each version runs on the bench through galois-edge only after an engineer approves it. Every step records its measured value, limits, pass or fail, raw command and response, instrument, operator, DUT serial and timestamps. Évariste compares the EVT and DVT runs, which you check against the recorded steps, and generates the report, while you choose the PVT steps and guard bands for the production sequence.
- Can you skip EVT or DVT?
- You can merge stages when the risk is low, such as a minor revision of a proven board, but each stage's question still needs an answer. Skipping DVT means the first proof that the design meets its requirements comes from production units, and any design change found then also changes tooling, fixtures and the production test.
- Is DVT the same as design verification?
- Often, but not by definition. Some companies expand DVT as design verification test, and stage names vary between companies. In regulated industries, design verification and design validation are defined terms with required records, so map the DVT plan onto those terms explicitly rather than assuming the build name covers them.
Bring Galois to your bench.
The daemon is Apache-2.0, free forever. Enterprise runs in your cloud or on-prem.