Skip to content
Safety Tech Review
Menu

Part 4: Buying and governance · Chapter 14

Buying and piloting safety technology

A practical buyer's guide to AI safety technology: business case, stakeholders, RFP questions, pilot design, acceptance criteria, contract terms and a checklist.

By · Updated · 19 min read · 12 sources · 1 figure

Buying safety technology well means starting from a defined hazard, building a business case that finance and workers can both accept, involving every function that can stop the project, and testing the product in a structured pilot with acceptance criteria agreed in advance. The most common failures are buying a capability before defining the problem, judging accuracy on demo footage, running pilots too short to show anything, and signing contracts that leave data ownership, model training and exit terms vague. This chapter walks through each stage and ends with a buyer checklist.

The advice applies to AI video analytics, wearables, proximity warning systems and EHS software, with most examples drawn from video analytics because it raises the widest set of questions.

Start with the problem

The strongest projects begin with a sentence like "We had 14 recorded vehicle and pedestrian near misses in the north yard last year and two lost-time injuries, and our controls depend on people following marked walkways." The weakest begin with "We should be using AI."

Before talking to vendors, write a one-page problem statement:

  • The hazard and where it occurs.
  • The evidence: incidents, near misses, audit findings, complaints, insurer recommendations.
  • Existing controls and why they are not enough.
  • Who is exposed, on which shifts.
  • What a good outcome looks like, in observable terms.

Then apply the hierarchy of controls. If the hazard can be removed or engineered out (physical segregation of pedestrians and vehicles, interlocks, speed limiters), that is usually more reliable than detection and alerting. Monitoring technology sits mostly in the administrative layer: it tells people something is wrong so they can act. It earns its place where engineering controls are impractical, where you need evidence about how work is really done, or where it can direct engineering fixes to the right places.

ISO 45001, the international standard for occupational health and safety management systems, expects organizations to control the risks introduced by procurement of products and services (clause 8.1.4) and to manage changes that affect health and safety performance (clause 8.1.3) [4]. A new monitoring system is both.

How do you build a business case for safety technology?

A business case that will survive a finance review has four parts: the cost of the current situation, the expected effect of the technology, the full cost of the technology, and the non-financial reasons that matter to the organization.

Cost of the status quo

Use your own numbers first. Useful inputs include:

  • Workers' compensation or employer liability claims costs for the hazard category over three to five years.
  • Lost time, restricted duty and overtime to cover absence.
  • Property and equipment damage from collisions and drops, which is often larger than injury cost and better recorded.
  • Downtime and investigation time after incidents.
  • Insurance premiums and any insurer conditions.
  • Regulatory exposure: past citations, improvement notices or enforcement.

National statistics give context but cannot replace your own figures. In the US, private-industry employers reported 2.5 million nonfatal injuries and illnesses in 2024, a rate of 2.3 cases per 100 full-time equivalent workers [1]. In Great Britain, HSE estimates 680,000 workers sustained a non-fatal injury at work in 2024/25, and that workplace injury and new cases of ill health cost £22.9 billion in 2023/24, the latest year for its cost estimate [2]. OSHA's business case page cites Liberty Mutual's 2025 estimate that US employers pay over $1 billion a week in direct workers' compensation costs for disabling, non-fatal injuries [3]. These figures tell a board that the problem is real. They do not tell you what your site will save.

Expected effect

This is where most business cases overreach. Vendors publish customer outcomes such as percentage reductions in unsafe events or incidents. Treat these as claims made by the vendor, usually from self-selected customers, without independent verification, without control groups and often measured in detected events rather than injuries. Chapter 15 explains why such figures are hard to interpret.

A defensible approach:

  • Model a range (low, central, high) rather than a single figure.
  • Base the central case on leading-indicator changes you can measure during a pilot, such as fewer pedestrian entries into vehicle zones or faster closure of hazard actions, and state the assumption that links them to injury reduction.
  • Include the cost of acting on the data: supervisor time to review alerts, engineering changes the data will prompt, training.
  • Run the case with a zero-injury-benefit scenario. If the investment still makes sense on damage reduction, insurance or regulatory grounds, say so.

OSHA's business case page points to a 2012 study of randomized Cal/OSHA inspections that found a 9.4% drop in injury claims and a 26% fall in workers' compensation costs over four years at inspected firms [3]. Well-designed evidence on safety interventions does exist, and effects of 10% to 25% are already substantial, so be skeptical of business cases that assume far larger reductions from software alone.

Total cost of ownership

Cost element What to include
Licenses Per camera, per site or per user fees; minimum terms; annual escalators
Hardware New cameras, edge compute devices, mounts, network switches, power
Installation Cabling, lifts and access equipment, contractor time, site shutdown if needed
Integration Connection to video management system, EHS software, identity and access management
Internal labor Project management, IT and OT time, alert review by supervisors, training
Governance DPIA, legal review, works council or union negotiation, audits
Change management Communications, briefings, refresher training
Ongoing Camera maintenance, re-calibration after layout changes, support tiers
Exit Data export, hardware removal or redeployment

Alert review time is the most commonly underestimated cost. A system producing 200 alerts per shift, each taking 30 seconds to review, consumes more than an hour and a half of someone's time per shift.

Non-financial reasons

Some benefits are real but hard to price: evidence for investigations, better visibility on night shifts, support for a site with a poor record, insurer expectations, or a group-wide commitment to a safety strategy. State them plainly rather than converting them into invented dollar figures.

Who needs to be involved?

Safety technology crosses more organizational boundaries than most EHS purchases. Each group below has a legitimate interest and, in many organizations, the power to stop or delay the project.

Stakeholder What they care about What to ask of them
EHS / HSE lead Hazard reduction, credibility of data, workload Own the problem statement, success measures and use policy
Operations and site managers Throughput, disruption, supervisor time Commit time to act on alerts; nominate pilot areas
Frontline supervisors Alert fatigue, conflict with crews Shape alert routing and review workflow
Workers and their representatives Surveillance, discipline, fairness Consult on purpose, limits and use policy; works agreement where required
IT Network, cybersecurity, identity, cloud Approve architecture and security controls
OT / engineering Camera placement, PLC or machine integration, maintenance Validate installation plan and integration
Physical security Existing CCTV, VMS, access control Clarify shared use of cameras and footage
Legal and data protection officer Lawful basis, DPIA, contracts, AI Act Lead DPIA and contract review
HR Disciplinary policy, employee relations Align disciplinary policy with the use policy
Procurement Competition, terms, supplier risk Run the RFP and negotiate
Finance Business case, total cost Agree assumptions before the pilot, not after
Insurer or broker Risk improvement Confirm whether results could affect terms

Hold a short kick-off with all of them before issuing the RFP. Ask each to write down what would make them say no. Those answers become requirements.

What should the requirements say?

Write requirements as outcomes and constraints, not as a feature list copied from a brochure. Group them:

  • Use cases. The specific hazards and detections required, ranked by priority, with the zones and camera views where they apply.
  • Performance. How you will measure precision, recall, latency and uptime (see acceptance criteria below).
  • Deployment. Use of existing cameras versus new, edge or cloud processing, network constraints, power, environmental ratings for outdoor or harsh areas.
  • Privacy by design. Anonymization options, ability to disable face recognition and identity features, configurable retention, regional data hosting.
  • Security. Certifications, encryption, access control, logging, vulnerability management, segregation from OT networks.
  • Integration. VMS, EHS incident and action management, single sign-on, data export formats and APIs.
  • Workflow. Alert routing, review and dismissal, action assignment and closure, reporting.
  • Support and service. Onboarding, training, response times, local language support, maintenance.
  • Commercial. Pricing model, term, pilot conversion, exit.

What questions should an RFP ask vendors?

The following questions are designed to separate capability from marketing, so ask for evidence with every answer.

Product and detection performance

  1. Which detection types are generally available today, which are in beta, and which are on the roadmap? Which are you proposing for our use cases?
  2. How do you measure precision and recall for each detection type? Provide recent results and describe the dataset (environments, lighting, camera angles, number of hours).
  3. How does performance change with camera height, angle, resolution, frame rate, night conditions, rain, dust, glare and occlusion?
  4. Have you tested for differences in detection rates across body types, skin tones, clothing, headwear and mobility aids? What did you find?
  5. What is a typical alert volume per camera per shift in a deployment like ours, and how is it tuned?
  6. How are site-specific zones, rules and exceptions configured, and by whom?
  7. How often are models updated, how are updates validated, and can we defer an update?

Architecture and deployment

  1. Where does inference run (on camera, edge device, on-premises server, cloud)? What leaves the site, in what form?
  2. What are the minimum camera specifications? Which of our existing cameras are suitable?
  3. What network bandwidth, power and rack space are required?
  4. How is the system monitored for camera outages, obstructions or drift in view?

Data protection and worker privacy

  1. What personal data is processed, at each stage? Do any features extract biometric identifiers such as face geometry?
  2. Can face blurring or full anonymization be applied before data leaves the device? Can identity features be disabled entirely?
  3. What are the default and configurable retention periods for clips, images and metadata?
  4. Do you use customer footage to train or improve models? If so, is it opt-in, how is it anonymized, and can we withdraw?
  5. Where is data hosted? List all sub-processors and their locations.
  6. Will you sign our data processing agreement under GDPR Article 28 [7]? Provide your standard terms.
  7. What documentation will you provide to support our DPIA [8]?

AI governance

  1. Do you consider any part of the product a high-risk AI system under the EU AI Act? What is your plan to meet provider obligations, and what instructions for use will you supply to deployers [9]?
  2. Does the product infer emotions, stress or engagement? (In the EU, workplace emotion recognition is prohibited outside narrow medical and safety exceptions; see Chapter 13.)
  3. Do you follow a recognized AI risk framework, such as the NIST AI Risk Management Framework [5] or ISO/IEC 42001?
  4. How do you handle reported errors, bias concerns or incidents caused by the system?

Security

  1. Which security certifications or attestations do you hold (for example ISO/IEC 27001 [10] or a SOC 2 Type II report)? Provide the latest report or certificate and its scope.
  2. How are edge devices hardened, patched and remotely managed? Who has remote access?
  3. How do you separate the system from operational technology networks?
  4. What is your incident notification commitment for security breaches?

Integration and workflow

  1. Which VMS and EHS platforms do you integrate with today, and how (native connector, API, file export)?
  2. Can alerts create actions in our EHS system, and can action status flow back?
  3. Can we export all event data in an open format at any time?

Commercial and company

  1. Describe your pricing model and every component of cost over three years, including hardware, installation, support and price escalators.
  2. What are the terms for a pilot and for converting a pilot to production?
  3. Provide three references with similar operations, at least one of which has been live for more than a year, and one customer that stopped using the product.
  4. What is your funding position or financial standing? What happens to our data and service if you are acquired or cease trading?

Score answers against weighted criteria agreed before the RFP goes out. Weight evidence of performance in conditions like yours, privacy design and total cost more heavily than feature breadth.

How do you design a pilot that proves something?

A pilot is an experiment. Treat it as one and it will answer the question "Should we scale this?" Treat it as a demo and it will answer "Does the product work in a sales sense?", which you already knew.

Define the question and scope

Pick one to three use cases that matter most, in one or two areas of one site. State the questions:

  • Does the system detect the target events accurately enough in our conditions?
  • Do supervisors and teams act on the alerts in time?
  • Do the target leading indicators improve compared with the baseline and comparison area?
  • What does it cost to run, in licenses and people's time?
  • How do workers and supervisors experience it?

Choose pilot areas carefully

Avoid choosing the area that had the worst month last quarter just because it had the worst month. Areas selected for unusually bad results tend to improve on their own, a statistical effect known as regression to the mean [11]. Choose areas based on exposure and hazard, and where possible pick a similar comparison area that does not get the technology, or stagger rollout between areas. Barnett and colleagues recommend control groups and multiple baseline measurements as the main design defenses against regression to the mean [11].

Measure a baseline first

Collect at least several weeks of baseline data before alerts go live. Two practical options:

  • Silent mode. Run the system with detections recorded but no alerts sent. This gives a baseline of detected events in the same terms the system will use later, and lets you calibrate zones and thresholds.
  • Manual observation. Run structured observations in the pilot and comparison areas before and during the pilot, using the same checklist, so you have a measure that does not depend on the vendor's detections.

Also extract the area's injury, near-miss and damage records for at least the previous one to three years.

Run in phases

Phase Typical duration Activities Exit criteria
Preparation 2 to 6 weeks DPIA, worker consultation, use policy, site survey, installation DPIA signed off; workers briefed; cameras commissioned
Calibration (silent mode) 2 to 4 weeks Zones and thresholds tuned; baseline collected; false positives reviewed Precision on reviewed alerts meets interim threshold
Active pilot 8 to 12 weeks or more Alerts live; review workflow; actions logged; weekly reviews Acceptance criteria measured
Evaluation 2 to 4 weeks Analysis, worker feedback, cost review, decision Go, extend or stop decision documented
Gantt chart of four pilot phases: preparation 2 to 6 weeks, calibration 2 to 4 weeks, active pilot 8 to 12 weeks or more, evaluation 2 to 4 weeks, for 14 to 26 weeks in total
Figure 14.1. The phases in the table above laid end to end. Solid bars show the shortest case and light bars the longest, so a full pilot takes roughly 14 to 26 weeks.

Durations are typical planning ranges, not rules. Rare events need longer pilots. Sites with strong seasonal patterns may need a pilot that spans a peak.

Decide in advance how alerts are used

During a pilot, alerts should drive coaching and engineering fixes, not discipline. Write this into the pilot use policy. If workers suspect the pilot is a disciplinary tool, behavior around cameras changes for the wrong reasons and near-miss reporting can drop, which corrupts the results. The ICO's monitoring guidance stresses telling workers what monitoring is for and keeping it proportionate [12].

Staff the pilot

Name an owner for each of: alert review, action follow-up, vendor liaison, data analysis and worker communication. Pilots without a named person reviewing alerts every shift fail quietly.

What acceptance criteria should a pilot use?

Write the criteria before go-live and attach them to the pilot contract. Each criterion needs a definition, a measurement method, a threshold and an owner.

Criterion Definition How to measure Example threshold (set your own)
Precision per detection type Share of alerts that are true events Reviewer labels a random sample of alerts each week Agreed per detection; higher for alerts that stop work
Recall per detection type Share of real events the system catches Staged or seeded scenarios; comparison with manual observation samples Agreed per detection and condition (day, night)
Latency Time from event to alert reaching the right person System logs Seconds for real-time interventions; minutes for reports
Uptime Share of scheduled time cameras and analytics are working System health logs Agreed service level
Alert volume Alerts per camera or area per shift System logs Low enough that every high-priority alert is reviewed
Time to action Time from alert or trend to assigned corrective action EHS action log Agreed target, such as within one shift for high severity
Action closure Share of actions closed on time EHS action log Agreed target
Leading indicator change Change in target behavior versus baseline and comparison area Silent-mode baseline, observations Direction and size agreed in advance
User experience Supervisor and worker views Short structured survey and interviews No unresolved serious concerns
Privacy and governance Configuration matches DPIA and use policy Audit of settings, access logs and any data requests Full conformance
Cost to operate Internal hours plus vendor fees Time logs, invoices Within business case range

Measure precision and recall separately. A system can achieve high precision by alerting rarely and missing many events, or high recall by alerting on everything. Vendors sometimes quote a single "accuracy" figure; ask what it means.

Staged tests are the most reliable way to measure recall. With worker consent and proper controls, walk a person through a marked exclusion zone, park a forklift in a pedestrian aisle, or remove a hard hat in view, across different times of day and camera positions. Never stage scenarios that create real risk.

What should the contract cover?

Many problems with safety technology surface at renewal or exit, when the buyer has the least leverage. Settle these points at signature.

Data and privacy terms

  • A data processing agreement meeting GDPR Article 28 where applicable, listing sub-processors and requiring notice of changes [7].
  • Customer ownership of footage, events and derived data, with the right to export at any time in an open format.
  • Clear terms on whether the vendor may use your data to train or improve models, defaulting to no unless you opt in, with anonymization requirements if you do.
  • Retention periods matching your DPIA, and certified deletion at the end of the contract.
  • Regional hosting commitments and international transfer safeguards.
  • Assistance with DPIAs, data subject requests and regulator inquiries.

AI-specific terms

The European Commission publishes model contractual clauses for public buyers procuring AI, in high-risk and non-high-risk versions; they cover areas such as risk management, data governance, transparency, human oversight and accuracy, and exclude commercial matters such as liability and payment [6]. Private buyers can borrow from them. Useful terms include:

  • A statement of intended purpose that matches your DPIA and use policy.
  • Vendor commitments on documentation, instructions for use and notification of material model changes.
  • Allocation of EU AI Act responsibilities if the system is or becomes high-risk, including support for your deployer duties such as informing workers' representatives [9].
  • A warranty that the product does not perform emotion recognition on workers or biometric identification unless explicitly specified.
  • Rights to performance information so you can monitor accuracy over time.

Security terms

  • Maintenance of named certifications or attestations [10], with the right to receive updated reports.
  • Patch timelines for critical vulnerabilities in edge devices.
  • Breach notification within a fixed number of hours.
  • Controls over vendor remote access.

Service and commercial terms

  • Service levels for uptime, support response and repair, with service credits.
  • Price protection: caps on annual increases and on per-camera fees as you scale.
  • Pilot conversion terms: how pilot fees are credited, and the price for the first production phase.
  • Hardware ownership and what happens to devices at exit.
  • Termination for convenience after an initial period, and termination rights if the vendor is acquired or changes the product materially.
  • Exit assistance: data export, deletion certificate, reasonable transition support.
  • Liability and indemnities, including for data protection breaches. Remember that safety technology does not transfer your legal duty of care to the vendor.

Common buying mistakes

  • Buying detections instead of outcomes. A long list of detection types matters less than three that address your main hazards and that people will act on.
  • Testing on the vendor's footage. Demo videos are chosen to work. Your night shift, rain and dust are what count.
  • No baseline. Without one, any change after go-live is uninterpretable.
  • Pilots that are too short. A few weeks shows alert volumes, not sustained behavior change.
  • Nobody owns alerts. Unreviewed alerts produce no safety benefit and do produce liability questions after an incident.
  • Workers hear about it last. Late consultation turns a safety project into an employee relations dispute.
  • Scaling before the workflow works. Fix alert routing and action closure at one site before adding ten.
  • Vague data terms. Ambiguity over model training and export is hard to fix later.

Buyer checklist

Stage Item Owner Done
Problem One-page problem statement with incident and near-miss evidence EHS ☐
Problem Hierarchy of controls reviewed; technology justified over engineering controls EHS, engineering ☐
Business case Cost of status quo from own data EHS, finance ☐
Business case Low, central and high benefit scenarios, including zero injury benefit Finance ☐
Business case Three-year total cost of ownership, including alert review time Finance, operations ☐
Stakeholders Kick-off held with all functions; each stated their conditions for "no" Project lead ☐
Workers Worker representatives consulted on purpose and limits; works agreement where required HR, EHS ☐
Requirements Use cases ranked; performance, privacy, security, integration requirements written EHS, IT ☐
RFP Questions issued; weighted scoring agreed before responses arrive Procurement ☐
RFP References checked, including a long-running customer and one that left Procurement ☐
Legal DPIA completed for the pilot DPO, legal ☐
Legal AI Act classification assessed and documented Legal ☐
Legal Data processing agreement and sub-processor list reviewed Legal ☐
Pilot Pilot and comparison areas chosen on exposure, not on last month's results EHS ☐
Pilot Baseline collected (silent mode and/or manual observation) EHS, vendor ☐
Pilot Use policy published; no discipline from pilot alerts HR, EHS ☐
Pilot Named owners for alert review and action follow-up Operations ☐
Acceptance Written criteria with definitions, methods and thresholds EHS, vendor ☐
Acceptance Precision sampled weekly; recall tested with safe staged scenarios EHS ☐
Contract Data ownership, export, training use, retention and deletion settled Legal, procurement ☐
Contract Security commitments, SLAs, price caps, exit terms agreed Procurement, IT ☐
Decision Evaluation report with results, costs and worker feedback; go, extend or stop Steering group ☐

Summary

Successful safety technology purchases start with a defined hazard and the evidence behind it, and they test whether monitoring is the right control before buying it. A credible business case uses the organization's own costs, models a range of effects rather than repeating vendor outcome claims, counts the full cost of ownership including the time it takes to review alerts, and states non-financial reasons honestly. Every function that can stop the project, from IT and the data protection officer to worker representatives, should shape the requirements before the RFP. RFP questions should demand evidence on accuracy in conditions like yours, privacy design, security, AI governance and total cost. A pilot should be run as an experiment, with areas chosen on exposure rather than recent bad results, a measured baseline, a comparison area where possible, a silent-mode calibration phase, alerts used for fixes rather than discipline, and written acceptance criteria for precision, recall, latency, uptime, alert volume, time to action and cost. Contracts should settle data ownership and export, model training on your data, retention, security, service levels, price protection, AI Act responsibilities and exit before you sign.

Frequently asked questions

+How long should a pilot of AI video safety analytics run?

Long enough to capture normal variation in shifts, seasons, staffing and volume, and to see whether actions follow alerts. For most sites that means at least 8 to 12 weeks after a calibration period, and longer if the target events are rare. Short pilots mostly measure alert volume, not safety outcomes.

+Should we pay for a pilot?

A modest paid pilot is often better than a free one. Payment gives you a contract with data protection terms, service commitments and clear ownership of results, and it makes both sides take the evaluation seriously. Agree in advance how pilot fees convert into a production contract if the pilot succeeds.

+What accuracy should we require?

There is no universal number. Set thresholds per detection type based on how alerts will be used: a high-severity alert that interrupts work needs very high precision, while a trend report can tolerate more noise. Measure precision from reviewed alerts and recall from staged or seeded events in your own environment.

+Can we run a pilot before finishing the DPIA?

In the EU and UK, a pilot that processes real workers' data is processing like any other, so a DPIA should be completed before it starts if the processing is likely to be high risk, which AI monitoring of workers usually is. A pilot DPIA can be narrower than the production one, and it gives you a head start on the full assessment.

Sources

  1. [1]US Bureau of Labor Statistics, Employer-reported workplace injuries and illnesses, 2024 (released January 2026)
  2. [2]HSE, Key figures for Great Britain 2024 to 2025
  3. [3]OSHA, Business case for safety and health
  4. [4]ISO 45001:2018 Occupational health and safety management systems
  5. [5]NIST, AI Risk Management Framework
  6. [6]European Commission Public Buyers Community, EU model contractual AI clauses
  7. [7]Regulation (EU) 2016/679 (General Data Protection Regulation), EUR-Lex
  8. [8]ICO, Data protection impact assessments
  9. [9]AI Act, Article 26: Obligations of deployers of high-risk AI systems
  10. [10]ISO/IEC 27001 Information security management systems
  11. [11]Barnett AG, van der Pols JC, Dobson AJ. Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology, 2005
  12. [12]ICO, Employment practices and data protection: monitoring workers

New chapters and updates, once a month

One email when we publish or update guidance. No vendor promotions. Unsubscribe any time.