Skip to content
Safety Tech Review
Menu

Part 1: History · Chapter 3

Computer Vision Comes to Work: From the Lab to the Shop Floor

How computer vision reached workplace safety: CCTV and IP cameras, ImageNet and AlexNet in 2012, deep learning, YOLO, edge AI hardware and the first AI safety startups.

By · Updated · 19 min read · 22 sources

Computer vision is the branch of artificial intelligence that lets software interpret images and video. It reached workplace safety through a chain of developments: decades of closed-circuit television (CCTV) for security, network cameras from the late 1990s, the deep learning breakthrough at the ImageNet competition in 2012, fast object detectors such as YOLO from 2015, and low-cost edge processors from around 2019. This chapter traces that chain, explains why each step mattered, and describes the first companies that applied the technology specifically to safety.

Why does the history of cameras matter to safety buyers?

Most AI video safety products sold today connect to cameras a site already has, which were usually installed for security, theft prevention or operational oversight. The capabilities and limits of that existing camera estate shape what a safety system can do. A camera mounted high over a loading dock to watch for intruders may be well placed to see forklift and pedestrian interactions but too far away to check whether someone is wearing safety glasses.

The history of the technology also explains why AI safety analytics arrived when it did, roughly in the late 2010s, and why earlier attempts at automated video monitoring disappointed. Understanding those limits helps buyers judge which claims are plausible today.

How did CCTV move from rocket tests to warehouses?

Closed-circuit television means cameras that send video to a limited set of monitors or recorders, rather than broadcasting it. One of the earliest recorded uses was in 1942, when Siemens installed a CCTV system at a test stand in Peenemünde, Germany, so engineers could observe V-2 rocket launches from a safe distance [1]. That first use was a safety application: keeping people away from a hazard while still letting them see it.

Commercial systems followed. The first commercial CCTV system in the United States became available in 1949 [1]. In September 1968 Olean, New York, became the first US city to install video cameras along its main business street for crime prevention [1]. In Britain, police in London installed cameras in the early 1960s, Bournemouth ran outdoor trials in 1985, and King's Lynn became the first local authority to use CCTV in 1987 [1]. By 2011 one estimate put the number of private and public cameras in the UK at about 1.85 million [1].

Recording and the analog era

Early CCTV was mostly watched live. The arrival of videocassette recorders in the 1970s made recording practical, and digital multiplexing in the 1990s let one recorder handle several cameras [1]. In industrial settings, cameras watched gates, yards, stores and production lines. Footage was reviewed after an incident, if it had been kept. Nobody expected a human to watch every screen in real time for safety hazards, and in practice nobody did.

This "record and review" pattern is the baseline that AI safety tools aim to replace. In the analog era, video was useful for investigations but almost never for prevention.

Network cameras and open standards

In 1996 the Swedish company Axis Communications launched what it describes as the world's first network camera, the AXIS Neteye 200 [2]. Network or IP (internet protocol) cameras send digital video over standard computer networks. That meant video could be stored on servers, accessed remotely and processed by software.

For years, though, cameras and recording software from different manufacturers often used incompatible protocols. In 2008 Axis, Bosch Security Systems and Sony founded ONVIF, an industry forum that publishes standard interfaces for IP-based security products [3]. ONVIF now has more than 500 member companies [3]. Standard interfaces made it much easier for a third-party analytics product to connect to a mix of cameras, which is how many AI safety vendors work today.

Period Camera technology How video was used for safety
1940s to 1960s Early analog CCTV, live viewing Watching hazardous tests from a distance
1970s to 1980s Analog cameras with VCR recording After-the-fact review of incidents and crime
1990s Digital multiplexing, first network camera (1996) Same, with easier storage
2000s Spread of IP cameras and video management software, ONVIF (2008) Remote viewing, simple motion detection
2010s High-definition IP cameras, cheaper storage, first deep learning analytics Early experiments with automated detection
2020s AI analytics on edge devices and in the cloud Continuous automated detection of safety events

What could computer vision do before deep learning?

Researchers have worked on computer vision since the 1960s. For most of that time, systems relied on features designed by hand: edges, corners, color histograms, texture patterns. Engineers wrote rules or trained simple classifiers on top of those features.

Viola-Jones and hand-built features

A well-known example is the face detector published by Paul Viola and Michael Jones in 2001. It used simple rectangular patterns called Haar features and a cascade of classifiers, and could detect faces in small images at about 15 frames per second on a 700 MHz desktop processor [4]. It was fast and compact enough to run on handheld devices of the time, but it was designed for frontal, upright, well-lit faces without occlusion [4].

That limitation is typical of the pre-deep learning era. Hand-built methods could work well on narrow tasks in controlled conditions. Industrial sites are rarely controlled. Lighting changes through the day, people wear bulky clothing and PPE, vehicles and loads block the view, cameras sit at awkward angles, and dust, rain and steam blur the image.

Video analytics in the 2000s

Security vendors in the 2000s sold video analytics such as motion detection, line crossing, loitering detection and left-object alerts. Most were based on detecting changes in pixels between frames. They could be useful for perimeter security at night, when any movement was suspicious. In a busy workplace, where people and vehicles move constantly and shadows shift, they produced large numbers of false alarms. Many security teams switched them off.

This experience left a lasting skepticism about automated video alerts that AI safety vendors still encounter. It also illustrates a point developed in Chapter 4: the value of an alert system depends heavily on its false positive rate, because people stop responding to systems that cry wolf.

What happened at ImageNet in 2012?

The change that made modern AI video safety possible came from academic research on image recognition.

Building ImageNet

In 2006 the computer scientist Fei-Fei Li began building ImageNet, a very large dataset of labeled images organized according to the structure of the WordNet lexical database [5]. The team first presented it as a poster at the 2009 Conference on Computer Vision and Pattern Recognition [5]. Labeling at that scale required crowdsourcing: between July 2008 and April 2010, about 49,000 workers from 167 countries filtered and labeled more than 160 million candidate images [5]. The full dataset grew to more than 14 million images across more than 20,000 categories [5].

From 2010 the project ran an annual competition, the ImageNet Large Scale Visual Recognition Challenge (ILSVRC), using a subset with 1,000 classes [5]. Teams were measured on top-5 error: the share of test images where the correct label was not among the model's five best guesses.

AlexNet

In 2012 a team from the University of Toronto, Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton, entered a deep convolutional neural network later known as AlexNet. It achieved a top-5 error of 15.3%, compared with 25.8% for the runner-up [6]. The network had about 60 million parameters across eight learned layers and was trained on two Nvidia GTX 580 graphics cards [6]. The AI researcher Yann LeCun later called it an unequivocal turning point in the history of computer vision [6].

Three ingredients came together in AlexNet:

  • Data: ImageNet supplied more than a million labeled training images.
  • Compute: graphics processing units (GPUs), designed for video games, could perform the massive parallel arithmetic that neural networks need.
  • Method: a deep network learned its own features from raw pixels instead of relying on hand-built ones.

The same three ingredients, data, compute and learned features, explain the strengths and weaknesses of AI safety systems today. They perform well on situations similar to their training data and less predictably on situations they have rarely seen.

Rapid progress after 2012

Progress was fast. In 2015 a Microsoft Research team introduced residual networks (ResNets), which made it possible to train networks with up to 152 layers. An ensemble of these networks achieved 3.57% top-5 error on the ImageNet test set and won ILSVRC 2015 [7]. By the final challenge in 2017, the winning error rate had fallen to about 2.25%, and the organizers concluded that the benchmark no longer posed a challenge [5].

Image classification answers "what is in this picture?" Safety applications need more. They need to know where each person, vehicle and object is, how they move, and how they relate to each other. That required progress in object detection.

How did object detection become fast enough for live video?

Object detection means finding each object in an image and drawing a box around it with a label, such as "person," "forklift" or "hard hat." It is the core task behind most video safety use cases: PPE checks, exclusion zones, vehicle and pedestrian proximity, and blocked exits.

Region-based detectors

Early deep learning detectors worked in two stages. They first proposed regions that might contain objects, then classified each region. Faster R-CNN, published in 2015 by Shaoqing Ren, Kaiming He, Ross Girshick and Jian Sun, made this process much more efficient by learning the region proposals inside the network. It ran at about 5 frames per second on a GPU with a large backbone network [8]. That was a big improvement, but still too slow for affordable analysis of many camera streams at once.

YOLO and single-stage detection

Also in 2015, Joseph Redmon, Santosh Divvala, Ross Girshick and Ali Farhadi published "You Only Look Once" (YOLO). Their approach treated detection as a single regression problem: one neural network looked at the whole image once and predicted boxes and classes directly [9]. The base model ran at 45 frames per second, and a smaller version reached 155 frames per second [9]. The authors noted that YOLO made more localization errors than leading detectors but was less likely to predict false detections in the background [9].

YOLO was followed by YOLOv2 in 2016 and YOLOv3 in 2018 [10]. Later versions, from YOLOv4 in 2020 onward, were developed by other researchers and companies not officially associated with the original authors [10]. The YOLO family, along with other single-stage detectors, became a common building block in commercial video analytics because it balanced speed and accuracy well. Buyers should note that different YOLO versions carry different software licenses, which can affect how vendors use them.

Datasets for everyday scenes

Detection also needed better training data. In 2014 a team led by Microsoft researchers released COCO (Common Objects in Context), with about 328,000 images, 91 object types and 2.5 million labeled instances, focused on objects in realistic scenes [11]. "Person" is one of the most heavily represented categories. Models pre-trained on COCO became the usual starting point for detecting people and vehicles, which vendors then fine-tune on their own industrial footage.

Pose estimation

A third capability, human pose estimation, finds the positions of body joints such as shoulders, elbows, hips and knees. In 2016 Zhe Cao and colleagues at Carnegie Mellon University published a method for real-time pose estimation of multiple people in an image, which placed first in the inaugural COCO 2016 keypoints challenge [12]. The same group later released it as part of the OpenPose software. Pose estimation underpins ergonomic analysis, fall detection and some behavior detection in safety products.

Year Milestone What it made possible
2001 Viola-Jones face detector Fast detection of one object type under controlled conditions
2009 ImageNet presented Large-scale training data
2012 AlexNet wins ILSVRC Deep learning beats hand-built features
2014 COCO dataset Training data for objects in real scenes
2015 Faster R-CNN, YOLO, ResNet Accurate detection; real-time speeds; very deep networks
2016 Real-time multi-person pose estimation Body posture analysis for ergonomics and falls
2019 $99 Nvidia Jetson Nano Affordable AI processing near the camera

Why did edge AI matter?

Running a detector on one video stream is easy on a powerful server. Running it on dozens or hundreds of streams around the clock is a cost and bandwidth problem. Sending all that high-definition video to the cloud can be expensive and may be impractical on sites with limited connectivity. It also raises privacy questions about where footage is stored.

Edge computing means processing data close to where it is generated, on a device at the site rather than in a distant data center. For video analytics, that can mean a small computer in a server room, a dedicated appliance, or processing inside the camera itself.

Nvidia's Jetson line illustrates how quickly edge hardware improved. The Jetson TK1 appeared in 2014 and the Jetson TX1 in 2015, each drawing about 10 watts [13]. In 2019 Nvidia released the Jetson Nano at $99, delivering about 0.47 teraflops of half-precision compute in 5 to 10 watts [13]. The Jetson Orin family followed in 2023, and the high-end Jetson AGX Thor in 2025 [13]. Camera makers also built more processing into cameras themselves; Axis, for example, designs its own ARTPEC chips for video processing and analysis [2].

Cheap edge hardware changed the economics. A site could run detection on many existing camera streams using a modest box on the premises, send only events and short clips to the cloud for review, and keep the bulk of raw footage local. Many AI safety products use some version of this hybrid design. Chapter 4 covers the tradeoffs between edge and cloud processing in detail.

Who were the first AI video safety companies?

Academic groups published research on vision-based safety monitoring for construction and industry through the 2010s, including work on detecting hard hats, tracking workers and equipment, and checking site progress. Commercial products dedicated to safety followed once the models and hardware described above were mature enough.

By the late 2010s and early 2020s, a group of startups had formed specifically to apply computer vision to workplace health and safety. Three examples illustrate the pattern. They are described here from their own published information, and their figures are company claims, not independently verified results.

Intenseye

Intenseye says it was founded in 2018 and is now headquartered in New York [14]. The company lists funding of a $4 million seed round in 2020, a $25 million Series A in 2021 and a $64 million Series B in 2024, with investors including Insight Partners, Lightspeed, Point Nine and Air Street Capital [14]. It describes its product as a real-time platform that analyzes video from existing CCTV or its own camera devices to detect more than 50 types of high-risk hazards, with a focus on potential serious injuries and fatalities [15]. Intenseye says it is live at more than 400 facilities in over 45 countries, and it publishes customer outcome figures, such as a 40% reduction in total recordable incident rate at the packaging manufacturer Huhtamaki [15]. These are vendor claims.

Voxel

Voxel, a San Francisco company, says it has raised $61 million and is deployed at more than 250 sites [17]. It describes its platform as analyzing existing camera infrastructure for safety and operational risk, and says it works with over 95% of existing IP cameras and adapts to new environments within 48 hours [18]. The company names customers including CEVA, MSI, Piston Automotive and the Port of Virginia, and says its customers report an average 65% reduction in workers' compensation claims [17] [18]. These are vendor claims.

Protex AI

Protex AI says it was founded in 2020 by Dan Hobbs and Ciaran O'Mara and has offices in Dublin and Boston [16]. It reports raising $54 million in total, with seed funding in 2021 and a Series A in 2022, and names investors including Hedosophia and Salesforce Ventures [16]. Protex describes its product as a platform that uses existing cameras to detect safety risks such as PPE non-compliance, ergonomic hazards and vehicle-pedestrian interactions, with integrations into EHS and other business systems. Its published case studies include claims such as an 80% reduction in incidents at Marks & Spencer [16]. These are vendor claims.

What these companies had in common

The early AI safety startups shared several design choices that still define the category:

  • They used existing IP cameras where possible to reduce installation cost and speed up deployment.
  • They processed video on site or in a hybrid edge and cloud design.
  • They focused on a library of specific, detectable events, such as missing PPE, people in vehicle routes, unsafe lifting postures, blocked exits and speeding vehicles.
  • They delivered events to dashboards and alerts for safety teams, often with short video clips for review.
  • They positioned their data as leading indicators, emphasizing prevention of serious injuries rather than counting minor ones.
  • They worked to address privacy concerns, for example by offering anonymization or face blurring.

Several larger players also moved in. Camera and video management vendors added safety analytics, EHS software companies began to integrate video events, and analyst firms started to track the space. Verdantix, for example, named computer vision among the developments expected to drive upgrades of safety management systems through 2028 [19]. Chapter 12 maps the current vendor categories and how they differ.

What problems did early deployments run into?

The first wave of commercial deployments surfaced issues that remain relevant to buyers.

Camera coverage rarely matched safety needs. Security cameras watch entrances, perimeters and valuable stock. Many safety hazards occur elsewhere: at racking, inside trailers, at machine interfaces or in yards with poor lighting. Effective deployments often needed some new cameras or repositioned ones.

Accuracy varied with conditions. Models trained mostly on daytime or indoor footage struggled with glare, night lighting, rain, fog and unusual viewpoints. Small objects such as safety glasses or gloves were harder to detect reliably from a distance than people or vehicles.

Alert volume could overwhelm teams. Detecting every instance of a minor rule breach in a busy facility can produce many events per shift. Without filtering and prioritization, safety teams faced the same alarm fatigue that undermined earlier video analytics.

Worker trust was not automatic. Continuous video analysis raises concerns about surveillance, discipline and privacy, especially where there is a history of poor labor relations. In Europe, the General Data Protection Regulation (GDPR), which has applied in all member states since 25 May 2018 [22], requires employers to justify and limit processing of personal data, including video. Chapter 13 covers the legal and ethical issues in depth.

The hierarchy of controls still applied. NIOSH ranks administrative controls and PPE below elimination, substitution and engineering controls because they depend on ongoing human effort [20]. Video detection of a PPE breach or a zone violation supports administrative controls. It does not remove the hazard. Deployments that produced real gains generally used the data to change layouts, traffic routes and processes, not only to correct individual workers.

Why does the timing matter now?

The safety problems that computer vision targets are persistent. In the US, the Bureau of Labor Statistics recorded 1,937 deaths from transportation incidents at work in 2024, 38% of all fatal work injuries, and pedestrian incidents rose 19% to 369 deaths [21]. Cameras can see many of these vehicle and pedestrian interactions in warehouses, yards, ports and construction sites.

The technology to address them only matured recently. The deep learning models, real-time detectors, training datasets and edge hardware described in this chapter all arrived between 2012 and 2019. The first dedicated companies formed around 2018 to 2020. That means most organizations evaluating AI video safety today are looking at a category less than a decade old, with limited independent evidence on long-term outcomes. Later chapters set out how to test claims, design pilots and measure results honestly. Newer methods, including vision-language models that can describe scenes in natural language, are covered as outlook in Chapter 16.

Summary

Computer vision reached workplace safety through a long chain of developments. CCTV began in 1942 with Siemens cameras for watching V-2 rocket tests and spread through security uses over the following decades. Axis launched the first network camera in 1996, and ONVIF, founded in 2008, standardized how cameras and software communicate, creating the installed base that AI safety tools now reuse.

Hand-built computer vision, such as the 2001 Viola-Jones face detector and the motion-based video analytics of the 2000s, worked in controlled conditions but produced too many errors in busy workplaces. ImageNet, presented in 2009, and AlexNet's win at ILSVRC 2012, with 15.3% top-5 error against 25.8% for the runner-up, began the deep learning era. ResNet reached 3.57% in 2015.

Object detection became fast enough for live video with Faster R-CNN and YOLO in 2015, with YOLO running at 45 frames per second. COCO supplied training data for real scenes, and multi-person pose estimation from 2016 enabled ergonomic analysis. Affordable edge hardware such as the $99 Nvidia Jetson Nano in 2019 made it practical to analyze many camera streams on site.

Dedicated AI video safety companies such as Intenseye, Voxel and Protex AI formed around 2018 to 2020, using existing cameras, edge processing and libraries of detectable safety events. Their published outcomes are vendor claims. Early deployments showed recurring challenges with camera coverage, accuracy in difficult conditions, alert volume, worker trust and the temptation to rely on monitoring instead of engineering out hazards.

Frequently asked questions

+Can AI safety software use the CCTV cameras we already have?

Often, yes. Most AI video safety products connect to existing IP cameras through standard video streams, and some vendors say most modern cameras are compatible. Camera position, resolution, lighting and network capacity still decide what can be detected, so expect a site survey and possibly some new cameras.

+Why did computer vision for safety only take off after 2015?

Three things had to arrive together: deep learning models that could recognize people and objects accurately (from 2012), detectors fast enough to run on live video (from about 2015), and affordable processors to run them near the cameras (from about 2019). Before that, systems were either too inaccurate or too expensive for routine use.

+Is computer vision for safety the same as facial recognition?

No. Most workplace safety systems detect people, vehicles, PPE, zones and postures without identifying individuals, and many vendors offer face blurring. Facial recognition is a separate capability with much stricter legal treatment in many jurisdictions, and buyers should confirm in writing whether it is used.

+What does YOLO mean in workplace safety?

YOLO (You Only Look Once) is a family of real-time object detection models first published in 2015. Many video analytics products, including safety tools, use YOLO-style detectors or similar single-stage models to find people, vehicles and equipment in each video frame.

Sources

  1. [1]Closed-circuit television (Wikipedia)
  2. [2]Axis Communications history (Axis)
  3. [3]ONVIF (Wikipedia)
  4. [4]Viola-Jones object detection framework (Wikipedia)
  5. [5]ImageNet (Wikipedia)
  6. [6]AlexNet (Wikipedia)
  7. [7]Deep Residual Learning for Image Recognition (He et al., arXiv 2015)
  8. [8]Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (Ren et al., arXiv 2015)
  9. [9]You Only Look Once: Unified, Real-Time Object Detection (Redmon et al., arXiv 2015)
  10. [10]You Only Look Once (Wikipedia)
  11. [11]Microsoft COCO: Common Objects in Context (Lin et al., arXiv 2014)
  12. [12]Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields (Cao et al., arXiv 2016)
  13. [13]Nvidia Jetson (Wikipedia)
  14. [14]Intenseye company page
  15. [15]Intenseye homepage
  16. [16]Protex AI company page
  17. [17]Voxel about page
  18. [18]Voxel homepage
  19. [19]EHS Software Market Size: The Road To $3 Billion (Verdantix)
  20. [20]Hierarchy of Controls (NIOSH)
  21. [21]Census of Fatal Occupational Injuries Summary, 2024 (BLS)
  22. [22]The general data protection regulation applies in all Member States from 25 May 2018 (EUR-Lex)

New chapters and updates, once a month

One email when we publish or update guidance. No vendor promotions. Unsubscribe any time.