Buying and piloting
AI safety pilot acceptance criteria: how to define pass and fail
How to write pass/fail acceptance criteria for an AI safety pilot: precision, recall, latency, alert load, time to action, sample sizes and a template.
By LIPAI WANG · Updated · 10 min read · 10 sources
Acceptance criteria are the written tests an AI safety pilot must pass before the buyer commits to a wider rollout. Good criteria are agreed before go-live, cover detection quality, operational fit, governance and cost, and state for each measure how it will be counted, what threshold counts as a pass and who is responsible. This article explains how to set them for AI video analytics and similar monitoring systems, with a template you can adapt.
The guidance applies to most safety monitoring technology, including proximity warning and wearables, but the examples come from video analytics, where accuracy claims are hardest to check. For the wider buying process, including RFPs and contracts, see the chapter on buying and piloting.
Why do pilots need written acceptance criteria?
Without agreed criteria, a pilot ends in a meeting where everyone reads the same dashboard differently. The vendor points to the number of unsafe events detected. Operations points to the supervisor hours spent clearing alerts. Finance asks what any of it means for injury costs. Each view is reasonable, and none of them answers whether the system should be scaled.
Written criteria fix three problems. They stop the goalposts moving after results arrive. They force a conversation, before money is spent, about what the system is for and how its alerts will be used. And they give procurement something to put in the contract, so that conversion to a paid production deployment depends on evidence rather than enthusiasm.
ISO 45001, the international standard for occupational health and safety management systems, expects organizations to manage changes that affect health and safety performance and to control risks from procured products and services [5]. A documented pilot with predefined acceptance tests is one practical way to show that a new monitoring system went through that process. The NIST AI Risk Management Framework makes a similar point from the AI side: its Measure function asks organizations to test AI systems against defined metrics in the conditions where they will be used [6].
What makes a good acceptance criterion?
Each criterion needs six parts. If one is missing, the criterion will be disputed at the end of the pilot.
| Part | What it means | Example |
|---|---|---|
| Definition | Exactly what is being counted | "A true alert is one where a reviewer confirms a person was inside the marked forklift exclusion zone while a forklift was within it" |
| Method | How the data is collected | Weekly random sample of 50 alerts, labeled by two trained reviewers |
| Threshold | The number that counts as a pass | At least 85% of sampled alerts are true |
| Sample size | How much data is needed before the result counts | At least 300 labeled alerts over the active phase |
| Conditions | Where and when the threshold must hold | Day and night shifts separately |
| Owner | Who measures and reports | Site EHS advisor, reviewed by the steering group |
It also helps to separate criteria into two kinds. Gates are must-pass conditions: if one fails, the pilot fails or is extended with a written fix plan. Scored measures inform the decision about how far and how fast to scale, and they can feed into the commercial negotiation. Privacy conformance and minimum recall on high-severity hazards are typical gates. User satisfaction and integration effort are typical scored measures.
How do you set precision and recall thresholds?
Precision is the share of alerts that are real events. Recall is the share of real events the system catches [1]. They pull against each other: tuning a model to alert more readily raises recall and lowers precision, and the reverse is also true [1]. A single "accuracy" figure hides this trade-off. Google's machine learning course gives the standard example: when the event of interest is rare, a model that never flags anything can still score 99% accuracy [1]. Unsafe events are rare compared with the hours of footage in which nothing happens, so ask any vendor quoting accuracy what the figure counts.
Set thresholds per detection type, according to what each alert triggers:
| Alert use | What matters most | Practical stance |
|---|---|---|
| Real-time intervention that stops work or sounds a horn | Precision | False alarms erode trust quickly; set a high precision gate |
| Real-time warning to a supervisor's phone | Both | Balance; cap alert volume per supervisor |
| Near-miss capture for investigation | Recall | Missed events defeat the purpose; accept more noise |
| Trend reporting and heat maps | Consistency | Errors matter less if they are stable over time and across areas |
There is no industry standard threshold. A figure you set should come from how alerts will be handled and how much reviewer time the site can afford. Agree interim thresholds for the calibration phase, when zones and sensitivity are still being tuned, and final thresholds for the active phase.
How many alerts do you need to review?
Precision measured on a handful of alerts is mostly noise. As a rough guide, if reviewers label 100 randomly sampled alerts and 90 are true, the 95% confidence interval runs from about 84% to 96%. With 400 labeled alerts at the same rate, it narrows to about 87% to 93%. That is ordinary sampling arithmetic, and it means a pilot target of "90% precision" needs several hundred reviewed alerts per detection type before anyone can say it was met.
Sample randomly across shifts, cameras and weeks. Reviewing only the alerts a supervisor happened to open will overrepresent busy periods and obvious events. Use two reviewers on part of the sample and compare their labels. If they disagree often, the definition of a true event needs tightening before the numbers mean anything.
How do you measure recall?
Recall is harder, because you need to know about the events the system missed. Three methods work in practice:
- Staged events. With worker agreement and proper controls, walk a person into a marked zone, park a vehicle in a pedestrian aisle or remove a hard hat in view, across different cameras, times of day and weather. Never stage anything that creates real risk.
- Independent observation. Trained observers record target events in a set area for a set period, using a checklist, and the results are matched against system detections.
- Footage review. Reviewers watch randomly selected footage segments and log every target event, then compare with what the system flagged.
Be careful with small numbers. Hanley and Lippman-Hand showed in 1983 that when no events occur in n trials, the upper 95% confidence limit for the true rate is roughly 3 divided by n [2]. Applied to recall testing, 30 staged events with no misses means the true miss rate could still be as high as about 10%. A recall gate on a high-severity hazard needs dozens of staged events per condition, spread across the cameras that matter.
Which operational criteria should a pilot include?
A system that detects well can still fail in operation. These criteria test whether the site can live with it.
| Criterion | Definition | Typical method |
|---|---|---|
| Latency | Time from event to alert reaching the right person | System logs, checked against staged events |
| Uptime | Share of scheduled hours each camera and the analytics are working | Health monitoring logs |
| Alert load | Alerts per supervisor per shift needing review | System logs |
| Review time | Minutes per shift spent reviewing alerts | Time logs or short diary study |
| Time to action | Time from alert or trend to an assigned corrective action | EHS action log |
| Action closure | Share of resulting actions closed by their due date | EHS action log |
| Running cost | License fees plus internal hours, scaled to full deployment | Invoices and time logs |
Alert load deserves a firm cap. If supervisors cannot review every high-priority alert during their shift, the system is generating a record of unaddressed hazards, which is a liability after an incident as well as a wasted cost. Time to action and closure tie the technology to the leading indicators OSHA describes as measures that drive change, as opposed to lagging indicators that count injuries after they happen [4].
How should leading indicator change be judged?
Most buyers want the pilot to show that unsafe behavior fell. That is a reasonable goal, but it needs a careful design.
First, collect a baseline. Running the system in silent mode for several weeks, with detections recorded but no alerts sent, gives a baseline measured in the same terms as the active phase. Structured manual observations in the same period give a second baseline that does not depend on the vendor.
Second, use a comparison area if you can. Areas picked because they had a bad recent month tend to improve on their own, an effect called regression to the mean, and Barnett and colleagues recommend comparison groups and repeated baseline measurements as the main defenses [3]. Without a comparison area, a fall in detected events could reflect the system, the extra attention a pilot brings, seasonal volume or chance.
Third, agree the direction and approximate size of change before go-live, and state which indicators count. Examples include pedestrian entries into vehicle zones per hundred vehicle movements, or the share of observed tasks with required equipment worn. Normalize by activity level so a quiet month does not look like an improvement.
Injury rates rarely belong in pilot acceptance criteria. Over a few months in one area, injuries are too infrequent to separate a real change from chance. Track them over the longer term instead.
Which governance criteria should be gates?
Governance failures can end a deployment regardless of performance, so treat these as gates:
- Privacy conformance. Settings such as face blurring, retention periods, access permissions and disabled features match the data protection impact assessment (DPIA). In the UK and EU, a DPIA is required before processing likely to result in high risk, and the ICO's guidance uses the tracking of employees as an example of processing that can meet that test [8]. Audit the configuration and the access logs at the end of the pilot.
- Use policy compliance. Alerts were used for coaching and engineering fixes, as agreed, and not for discipline outside the agreed policy. The ICO's monitoring guidance expects workers to be told what monitoring is for [9].
- Worker information. If the system is or may become a high-risk AI system under the EU AI Act, employers acting as deployers must inform workers' representatives and affected workers before putting it into service, and must assign human oversight to people with the competence and authority to exercise it [7]. Even where the Act does not apply, both are good practice for a pilot.
- Human oversight in practice. Named people reviewed alerts on every pilot shift, and their decisions were logged.
What does an acceptance criteria template look like?
Adapt the table below for each detection type in scope. The thresholds shown are placeholders for illustration; set your own based on alert use, staffing and risk.
| # | Criterion | Type | Threshold (example only) | Sample or period | Owner |
|---|---|---|---|---|---|
| 1 | Precision, vehicle zone intrusion | Gate | Agreed % of sampled alerts true, day and night separately | 300+ labeled alerts | EHS advisor |
| 2 | Recall, vehicle zone intrusion | Gate | Agreed % of staged events detected | 60+ staged events across cameras and shifts | EHS advisor |
| 3 | Latency, real-time alerts | Gate | Alert received within agreed seconds | All staged events | IT |
| 4 | Uptime | Gate | Agreed % of scheduled hours | Full active phase | IT, vendor |
| 5 | Alert load | Gate | Below agreed alerts per supervisor per shift | Full active phase | Operations |
| 6 | Time to action, high severity | Scored | Action assigned within one shift | All high-severity alerts | Operations |
| 7 | Leading indicator change | Scored | Agreed direction and size versus baseline and comparison area | Baseline plus active phase | EHS lead |
| 8 | Privacy and use policy conformance | Gate | Full conformance with DPIA and policy | End-of-pilot audit | Data protection lead |
| 9 | Supervisor and worker experience | Scored | No unresolved serious concerns | Survey and interviews | HR, EHS |
| 10 | Projected running cost | Scored | Within business case range | Invoices, time logs | Finance |
Attach the completed table to the pilot agreement. The European Commission's model contractual AI clauses for public buyers include provisions on accuracy, transparency and human oversight that private buyers can borrow when drafting the performance and documentation obligations around a pilot [10].
How do you make the go or no-go decision?
At the end of the active phase, produce a short evaluation report with one line per criterion: result, sample size, pass or fail, and comments. Then apply the rules agreed at the start:
- All gates passed: proceed to a scale-up plan, using scored measures to set scope and pace.
- A gate failed for a fixable reason, such as a poorly placed camera: extend the pilot for a fixed period with a written fix plan and the same criteria.
- A gate failed for a structural reason, such as recall that cannot be raised without unmanageable alert volume: stop, or scale only the detection types that passed.
Record the decision and the reasons. If the system goes into production, keep measuring the same criteria at intervals, because model updates, layout changes and seasonal conditions can shift performance after the pilot ends.
Summary
Acceptance criteria turn an AI safety pilot from a product demonstration into a test with a clear answer. Write them before go-live, attach them to the contract and give each one a definition, method, threshold, sample size, conditions and owner. Separate must-pass gates from scored measures. Set precision and recall per detection type according to what each alert triggers, and gather enough reviewed alerts and staged events for the results to be meaningful. Add operational criteria for latency, uptime, alert load, time to action and cost, judge leading indicator change against a baseline and a comparison area, and treat privacy conformance, use policy compliance, worker information and human oversight as gates. At the end, apply the decision rules you agreed at the start and keep measuring after rollout.
Frequently asked questions
+Who should sign off the acceptance criteria?
The EHS lead should own them, with sign-off from operations, IT, the data protection lead and finance, and agreement from the vendor. Worker representatives should see them before go-live. If any of these groups first sees the criteria at the end of the pilot, expect the result to be contested.
+Can the vendor measure the results for us?
The vendor can supply system logs, but labeling alerts as true or false and running staged tests should be done or supervised by your own people. Precision measured by the party that benefits from a high number is a vendor claim, not an independent result.
+What if the pilot passes some criteria and fails others?
Decide in advance which criteria are gates and which are scored. A failed gate means stop or extend with a written fix plan. Mixed results on scored measures feed the scale-up decision and the commercial negotiation, for example by limiting the rollout to detection types that passed.
+Should injury rates be an acceptance criterion?
Usually not for a pilot. Injuries in a single area over a few months are too rare for a change to be distinguished from chance. Use leading indicators and operational measures for the pass or fail decision, and track injury and damage records over the longer term.
Related reading
Sources
- [1]Google for Developers, Classification: accuracy, recall, precision and related metrics
- [2]Hanley JA, Lippman-Hand A. If nothing goes wrong, is everything all right? Interpreting zero numerators. JAMA, 1983
- [3]Barnett AG, van der Pols JC, Dobson AJ. Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology, 2005
- [4]OSHA, Using leading indicators to improve safety and health outcomes
- [5]ISO 45001:2018 Occupational health and safety management systems
- [6]NIST, AI Risk Management Framework
- [7]AI Act, Article 26: Obligations of deployers of high-risk AI systems
- [8]ICO, When do we need to do a DPIA?
- [9]ICO, Employment practices and data protection: monitoring workers
- [10]European Commission Public Buyers Community, EU model contractual AI clauses
New chapters and updates, once a month
One email when we publish or update guidance. No vendor promotions. Unsubscribe any time.