Skip to content
Safety Tech Review
Menu

Buying and piloting

40 questions to ask AI safety vendors: an RFP checklist

A 40-question RFP checklist for AI safety vendors, grouped by fit, accuracy, deployment, privacy, AI governance, security, workflow and commercial terms.

By · Updated · 11 min read · 14 sources

This checklist gives 40 questions to put to vendors of AI safety technology, such as video analytics that detect unsafe acts and conditions, grouped into eight areas: fit, detection performance, deployment, data protection, AI governance, security, workflow and commercial terms. Each question comes with what a good answer looks like and the warning signs to watch for. Use it to build a request for proposal (RFP) or a shorter request for information, and to compare answers on evidence rather than presentation.

The questions assume a buyer in the US, UK or EU evaluating AI video analytics. Most of them also work for proximity warning systems, wearables and EHS software with AI features.

How should you use this checklist?

Before sending any questions, write down the hazard you are trying to control, where it occurs and what evidence you have, such as incident and near-miss records. Vendors answer better when they know the problem, and you can then score answers against it.

Three practical rules make the answers easier to compare:

  • Ask for documents. Test reports, certificates, sample contracts and data processing agreements are harder to inflate than prose answers.
  • Fix the format. Give a word limit per answer and a table for pricing so responses line up.
  • Agree scoring first. Set weights for each group before responses arrive, and have at least two people score independently.

Whatever a vendor says about outcomes at other customers, such as reductions in incidents, remains the vendor's claim. Record it as such and test what you can in a pilot.

Group 1: Fit and use cases (questions 1 to 5)

These questions check whether the product addresses your hazard or only a nearby one.

# Question A good answer includes Red flags
1 Which of our listed hazards can your product detect today, and which detections are in beta or on the roadmap? A clear split between generally available, beta and planned, mapped to your hazards Roadmap items presented as available
2 Which customers use these detections in operations similar to ours (sector, site size, indoor or outdoor)? Named or anonymized examples with comparable conditions Only demo footage or unrelated sectors
3 What does your product not do well, and in which conditions do you advise against it? Specific limits, such as long distances, heavy occlusion or low light "No limitations"
4 How are site-specific rules, zones and exceptions configured, and who does it? A described process, with customer and vendor roles Every change needs a paid vendor request
5 What changes on our side (layout, procedures, staffing) have customers needed to make to get value? Honest examples of supervisor time and process changes "Plug and play" with no workload

Group 2: Detection performance and evidence (questions 6 to 11)

Precision is the share of alerts that are real events; recall is the share of real events the system catches [13]. Improving one usually lowers the other, and a single accuracy figure can look high for a system that misses rare events [13]. Ask for both, per detection type.

# Question A good answer includes Red flags
6 What are precision and recall for each proposed detection, and on what test data (hours, sites, lighting, camera angles)? Figures per detection with a described dataset and date One headline "accuracy" number
7 How does performance change at night, in rain, dust, glare, at long range or with partial occlusion? Results broken down by condition, or a plan to test them in your pilot No condition-level data
8 Have you tested whether detection rates differ by body size, skin tone, clothing, headwear, PPE color or mobility aids? What did you find? A described test and its results, including any gaps found "Our AI has no bias"
9 What alert volume per camera per shift should we expect, and how is it tuned down without losing important events? A typical range from comparable sites and a tuning method No figures
10 Will you support a pilot measured on our own footage, with our reviewers labeling alerts and staged tests for recall? Agreement, plus help with test design Results reported only from the vendor's dashboard
11 How often do models change, how is each update validated, and can we defer or roll back an update? Release notes, regression testing and customer control Silent updates with no notice

Question 8 matters because demographic differences in error rates are well documented in some computer vision tasks. A 2019 NIST study of 189 face recognition algorithms found demographic differentials in false positive rates in the majority of them [12]. Safety detections such as person or PPE detection are different tasks, but the study shows why vendors should test for such gaps instead of assuming they are absent.

Group 3: Deployment and architecture (questions 12 to 16)

# Question A good answer includes Red flags
12 Where does processing run (camera, edge device, on-premises server, cloud), and what data leaves the site, in what form? A data flow diagram showing video, clips, images and metadata Vague "secure cloud" answer
13 What are the minimum camera specifications, and which of our existing cameras qualify? A site survey method and clear specifications All cameras qualify without a survey
14 What bandwidth, power, rack space and network changes are required? Figures per camera and per site Requirements discovered after contract
15 How do you detect camera outages, blocked views, moved cameras or degraded image quality? Automatic health monitoring and alerts Customer must notice problems
16 What happens to detection and alerting if the internet connection or cloud service is down? Local buffering or edge alerting, and stated recovery behavior Total loss of alerts with no notice

Group 4: Data protection and worker privacy (questions 17 to 22)

In the UK and EU, workplace video analytics usually requires a data protection impact assessment (DPIA) before processing starts, and the ICO's guidance uses employee tracking as an example of processing that can require one [2]. Under the GDPR, biometric data used to uniquely identify a person is a special category with stricter conditions, and vendors acting as processors must sign terms meeting Article 28 [1].

# Question A good answer includes Red flags
17 What personal data does the system process at each stage, and does any feature use biometric identification such as face matching? A data inventory per feature, with identity features clearly marked Unclear whether faces are matched
18 Can faces and bodies be blurred or anonymized before data leaves the device, and can identity features be switched off entirely? Configurable anonymization at the edge and a way to verify it Anonymization only in the user interface
19 Do you use customer footage to train or improve models? Is it opt-in, how is it de-identified, and can we withdraw? Off by default, written opt-in, described de-identification, withdrawal process Training rights buried in terms of service
20 What are the default and configurable retention periods for video, clips, images and event data? Per data type, configurable by the customer One fixed retention period
21 Where is data hosted, which sub-processors are used, and will you sign our data processing agreement? Region options, a current sub-processor list, Article 28 terms [1] Refusal to share sub-processors
22 What documentation will you provide for our DPIA and worker consultation? A DPIA support pack and plain-language worker materials "That is the customer's job"

The ICO's monitoring guidance expects employers to be open with workers about what monitoring does and why [3]. Vendor materials written for workers, as asked in question 22, help with that.

Group 5: AI governance and regulation (questions 23 to 27)

The EU AI Act already prohibits AI systems that infer the emotions of people at work, except for medical or safety reasons; that prohibition has applied since 2 February 2025 [4][6]. Where a system is high-risk under the Act, employers deploying it must inform workers' representatives and affected workers before putting it into service and assign competent human oversight [5]. Under the amended application dates in Article 113, obligations for high-risk systems listed in Annex III apply from 2 December 2027 [6]. Ask now, because contracts signed today will run into that period.

# Question A good answer includes Red flags
23 Have you assessed whether any part of the product is a high-risk AI system under the EU AI Act? What is the conclusion and reasoning? A written assessment per feature "The AI Act does not apply to us" with no reasoning
24 Does any feature infer emotions, stress, fatigue or engagement from faces, voice or body language? A clear yes or no per feature, with the safety rationale for any exception claimed Fatigue or attention scoring with no legal analysis
25 What instructions for use, intended purpose statement and logging will you provide to support our deployer duties? Documents available now, plus log export Promised "when required"
26 Do you follow a recognized AI governance framework, such as the NIST AI Risk Management Framework or ISO/IEC 42001? Named framework and evidence of use, such as a certificate or internal policy [7][8] Framework named with no evidence
27 How do customers report errors or bias concerns, and how do you investigate and tell affected customers? A defined process with response times No process

Group 6: Security (questions 28 to 32)

# Question A good answer includes Red flags
28 Which security certifications or attestations do you hold, and what is their scope? Current ISO/IEC 27001 certificate or SOC 2 report, with scope covering the product and hosting [9][10] Certification of a data center only
29 How are edge devices hardened, patched and monitored, and how fast are critical vulnerabilities fixed? Stated patch timelines and secure boot or equivalent Manual patching on request
30 Who at your company can access our systems or footage remotely, and how is that access approved and logged? Named roles, customer approval, full logging Standing remote access for support staff
31 How is the system separated from our operational technology networks, and do you follow ISA/IEC 62443 guidance? Network segmentation design and familiarity with industrial security standards [11] Request for flat network access
32 How quickly will you notify us of a security incident affecting our data? A fixed number of hours in the contract "Promptly"

SOC 2 reports cover controls relevant to security, availability, processing integrity, confidentiality and privacy [10]. Check the report's scope and any exceptions the auditor noted.

Group 7: Workflow, integration and support (questions 33 to 36)

# Question A good answer includes Red flags
33 How are alerts routed, reviewed, dismissed and escalated, and can we cap alert volume per person? Configurable routing and caps, with review logged All alerts to everyone
34 Which video management and EHS systems do you integrate with, and can alerts create actions whose status flows back? Named integrations and a documented API Integration "on the roadmap"
35 Can we export all event data and clips in an open format at any time, at no extra cost? Self-service export in standard formats Export only at exit, for a fee
36 What onboarding, training and support do you provide, in which languages and time zones? Named service tiers and response times Support by email only

Group 8: Commercial terms, references and exit (questions 37 to 40)

# Question A good answer includes Red flags
37 What is the full three-year cost, including licenses, hardware, installation, support and annual price increases? A complete pricing table with caps on increases Hardware or support priced later
38 What are the pilot terms, and how do pilot fees convert into a production contract? Written pilot terms, acceptance criteria and conversion pricing Free pilot with no terms
39 Can you give three references with similar operations, one live for more than a year, and one customer that stopped using the product? All three, including the former customer Only recent or hand-picked references
40 What happens to our data and service if we leave, or if you are acquired or stop trading? Export, certified deletion, transition support and change-of-control terms No exit provisions

The European Commission's model contractual AI clauses, published in high-risk and non-high-risk versions for public buyers, are a useful source of contract wording on transparency, human oversight and accuracy [14]. They leave out liability and payment terms, so pair them with your standard commercial terms.

How should answers be scored?

A simple scheme works well:

Score Meaning
0 No answer, refusal or irrelevant answer
1 Answer given with no supporting evidence
2 Clear answer with partial evidence
3 Clear answer backed by documents, test results or contract terms

Weight the groups to match your priorities. For a first deployment of video analytics, many buyers give the most weight to detection performance, data protection and three-year cost, and less to breadth of features. Treat some answers as pass or fail regardless of score, for example a refusal to switch off identity features or to put a breach notification period in the contract.

Shortlist two or three vendors and test them in a structured pilot with written acceptance criteria. The RFP tells you what a vendor says it can do. The pilot tells you whether it does so on your site.

Summary

A good RFP for AI safety technology asks for evidence in every area: fit with your hazards, precision and recall per detection in conditions like yours, deployment and resilience, data protection and worker privacy, AI governance and EU AI Act readiness, security, workflow and integration, and commercial and exit terms. The 40 questions here, each with what a good answer includes and the red flags to watch for, can be used in full for a formal tender or cut down to the most important for a smaller request. Agree weighted scoring before answers arrive, treat vendor outcome figures as claims, ask for a reference that left, and confirm the shortlist through a pilot measured on your own footage.

Frequently asked questions

+Do we need to ask all 40 questions?

No. Use the full list for a formal RFP covering several sites, and pick the ten to fifteen most relevant for a smaller request for information. The questions on detection evidence, data use for model training, emotion recognition, security attestations and exit terms are worth keeping in every version.

+How should we score vendor answers?

Agree weights for each group before issuing the RFP, then score each answer on a simple scale such as 0 to 3, where 0 means no answer or a refusal and 3 means a clear answer backed by documents. Score the evidence offered, and have at least two people score independently.

+What if a vendor says the information is confidential?

Offer a non-disclosure agreement for detailed test results, security reports and contracts. A vendor that will not share a SOC 2 report or test methodology under NDA is asking you to accept its claims on trust, and that should cost it points.

+Should the same questions apply to proximity warning and wearable vendors?

Most of them apply with small changes. Replace the camera questions with questions about tags, batteries, radio coverage and charging, and keep the questions on detection performance, data, security, workflow and exit.

Sources

  1. [1]Regulation (EU) 2016/679 (General Data Protection Regulation), EUR-Lex
  2. [2]ICO, When do we need to do a DPIA?
  3. [3]ICO, Employment practices and data protection: monitoring workers
  4. [4]AI Act, Article 5: Prohibited AI practices
  5. [5]AI Act, Article 26: Obligations of deployers of high-risk AI systems
  6. [6]AI Act, Article 113: Entry into force and application
  7. [7]NIST, AI Risk Management Framework
  8. [8]ISO/IEC 42001 Information technology, Artificial intelligence, Management system
  9. [9]ISO/IEC 27001 Information security management systems
  10. [10]AICPA, SOC 2 examinations
  11. [11]ISA, ISA/IEC 62443 series of standards
  12. [12]NIST, NIST study evaluates effects of race, age, sex on face recognition software (2019)
  13. [13]Google for Developers, Classification: accuracy, recall, precision and related metrics
  14. [14]European Commission Public Buyers Community, EU model contractual AI clauses

New chapters and updates, once a month

One email when we publish or update guidance. No vendor promotions. Unsubscribe any time.