Site Search

AI Application Examples for Surveillance Cameras: Explaining the Mechanism of Real-Time Video Analysis Achieved with YOLO x VLM

Surveillance cameras are moving from "recording" to "understanding the images."

In recent years, surveillance cameras have been used in a variety of locations, including factories, logistics facilities, commercial facilities, and public facilities, for purposes such as crime prevention, safety management, and equipment monitoring.

Furthermore, nowadays, the use of AI is expanding beyond simply recording video to include real-time analysis of camera footage.

 

In fact, the total domestic market size for surveillance camera systems in fiscal year 2024 is projected to be 225.4 billion yen, a 112.8% increase compared to the previous year, indicating an expanding trend (*).

One of the factors driving market expansion is AI-powered image analysis. The survey indicates the proliferation of edge AI cameras equipped with AI analysis capabilities and the growing use of real-time analysis that processes video immediately on-site.

*Source: Yano Research Institute, "Survey on the Domestic Market for Surveillance Cameras/Systems (2025)"

 

This type of AI video analysis utilizes object detection models such as YOLO to automatically detect people, vehicles, and specific objects from camera footage.

This allows for applications such as AI detecting in real time which scenes need to be checked, rather than humans constantly monitoring all the footage.

 

On the other hand, while object detection can determine "what is being shown," it can be difficult to determine the context in which that object exists.

For example, even if the same object is detected, the situation that needs to be checked will differ depending on whether a person is possessing the object and what other people and objects are present in the surrounding area. Therefore, it is important to understand the video not only by detecting the object, but also by including the surrounding environment.

 

Therefore, what we will introduce today is video analysis that combines object detection with generative AI such as VLM/LMM.

By detecting objects in real time to identify scenes that need to be checked, and then analyzing that footage with generating AI, we support AI not only in "detection" but also in "situational understanding."

 

This article will introduce the mechanism and potential applications of the AI platform SiMa.ai, based on a demo running on the platform.

How will AI-powered object detection and generation change video analysis?

Conventional object detection systems can detect people, vehicles, and specific objects from camera footage, allowing for real-time understanding of "what is where."

On the other hand, generative AI such as VLM (Vision Language Model) and LMM (Large Multimodal Model) can process images in combination with natural language, allowing them to describe not only the people and objects in a video, but also their relationships and the surrounding environment in natural language.

 

By combining these, you can use object detection as a starting point and then further analyze the video requiring verification using VLM/LMM.
The process can be summarized as follows:

 

① Analyze camera footage in real time using an object detection model.
② Detect the target object
③ Analyze the detected frames using VLM/LMM
④ Output the situation in the video in natural language.

 

In other words, object detection plays the role of "finding the scene that needs to be checked," while VLM/LMM plays the role of "understanding what is happening in that place."

The key point of this video analysis is that it can not only detect the presence or absence of objects, but also use AI to understand the relationships between people and objects, as well as the surrounding environment.

Video analysis using object detection and generation AI, as seen in the demo video.

Now, let's take a look at a video analysis demo running on the SiMa.ai AI platform. (This video is in English.)

In this demonstration, as an example of video analysis combining object detection and generative AI, we are using firearms as the target for detection, assuming the detection of dangerous objects.

You can see the entire process, from detecting objects in real time from camera footage to analyzing the situation on the scene using generated AI.

Now, let's take a look at the processes that are being performed in the video, step by step.

① Real-time detection of objects from multiple camera feeds

First, the video input from the camera is continuously analyzed using an object detection model such as YOLO.

In the demo, when a firearm appears in the video, the AI detects the target in real time and displays the detection results on the screen. This video analysis can be processed in parallel at up to 16 channels, VGA resolution, and 30fps.

By simultaneously analyzing footage from multiple cameras at the edge, objects can be detected on the spot without waiting for processing in the cloud. Even in environments with many cameras, AI can continuously analyze the footage and efficiently narrow down the scenes that require human review.

② The detected scene is analyzed and verbalized using a generative AI.

Next, the video footage after detecting the object is further analyzed using a generating AI.

In the demo, frames in which firearms are detected are input into the Gemma model, which is then instructed to describe the content of the images.

The system not only recognizes firearms, but also analyzes the number of people in the video, their clothing, and the objects they are carrying, and outputs this information in natural language.

 

Object detection alone primarily provides information such as "firearms are present," but by combining VLM/LMM, it becomes possible to understand who is carrying what, and what the surrounding environment is like.

This allows the AI not only to detect situations that need attention, but also to assist in understanding the subsequent situation.

③ Analyze the "context of the video" from the recorded footage.

Furthermore, this type of analysis using generative AI can be applied not only to real-time video but also to recorded video.

In the latter half of the demonstration, another video showing multiple people handling firearms is input into the VLM/LMM, and the instruction is to "explain it in 50 words or less."

The AI then not only recognizes that "there is a person with a firearm," but also describes the situation, including that multiple people are conducting firearms training outdoors, with an instructor guiding the participants.

 

The key point here is that even if the same object is shown, the meaning of the image can differ depending on the context.

While object detection captures "what is in the image," combining VLM/LMM allows for analysis of the context of the video, including the relationships between people and objects, and "what is happening there."

Furthermore, by having AI analyze recorded video footage as needed, or by having a program automatically analyze it, it is possible to apply this to reviewing and investigating past footage.

Examples of potential applications for object detection and generation AI

While this demonstration focuses on detecting firearms, the key point of this system is not the firearms themselves.

The mechanism of "detecting specific targets in real time and then using generating AI to further analyze the necessary scenes" can be applied to various situations in Japan by changing the detection targets and analysis content.

◎ Safety management and work status checks at the manufacturing site

In manufacturing environments, workers, forklifts, and transport robots often operate in the same space.

For example, by detecting people and vehicles in real time using object detection and analyzing situations requiring verification with VLM/LMM, it is possible to apply this to safety checks that include the positional relationship between workers and vehicles, as well as the surrounding environment.

Furthermore, by detecting protective equipment such as helmets and safety vests, as well as people entering restricted areas, it can also be used to verify safety rules.

The advantage of combining generative AI is that it can not only detect "people," "vehicles," and "protective equipment" individually, but also analyze their relationships and the surrounding environment.

◎ Understanding the status of people, vehicles, and cargo in logistics warehouses

In a logistics warehouse, many things move around in the same space, including workers, forklifts, AMRs, and goods.

By recognizing each object in real time through object detection and analyzing the situation requiring confirmation using VLM/LMM, it is possible to check situations such as when people and vehicles are approaching each other, or when luggage is placed in an unusual location.

By analyzing the positional relationships of multiple objects and the surrounding environment, this technology can help clarify situations that are difficult to judge using simple object detection alone.

 

*The mechanisms and application examples of AMRs are explained in more detail in another article.
Related article: Autonomous Mobile Robot/AMR Application Examples - Physical AI Realized with Visual SLAM and YOLOv8 (with demo video)

◎Safety management and situation checks at public and commercial facilities

In facilities with numerous cameras, such as train stations, shopping malls, and offices, it is not easy for a person to constantly monitor all the footage.

For example, by detecting specific objects such as abandoned luggage and analyzing the video footage with VLM/LMM, it's possible to determine information such as whether there are people nearby and the circumstances under which the object is placed.

Furthermore, when a person or specific object is detected, the AI generates verbal descriptions of the number of people present and the surrounding environment, which can assist surveillance personnel in understanding the situation.

Summary: From "detecting" objects to "understanding" situations through video analysis.

The demo we presented today combines real-time detection using object detection models such as YOLO with situational awareness of the video using VLM/LMM.

Object detection helps identify scenes that need to be examined, and then the resulting video is analyzed using AI to detect what is being shown and understand what is happening in that scene—a complete video analysis process.

Furthermore, this demo features parallel processing of up to 16 channels of video at VGA resolution and 30fps, enabling real-time analysis of multiple camera feeds at the edge.

 

SiMa.ai's AI platform enables various AI processes to be executed in an edge environment, from real-time object detection using YOLO and other technologies to video situation analysis using VLM/LMM.

SiMa.ai's MLSoC™ achieves high-efficiency inference of up to 50 TOPS while being small and low-power. Its key feature is that it can process everything from video acquisition and object detection to situational analysis using generative AI on the edge side, without relying on communication with the cloud.

  

 

We hope you will find this article useful as your first step in utilizing AI.

Inquiry

Please feel free to contact us with any questions about our products, technical inquiries, sample requests, or estimates.

SiMa.ai Manufacturer Information Top

Edge AI Use Cases (with Demo Videos)

For those looking for ideas for implementing edge AI, we are showcasing a demo utilizing SiMa.ai's MLSoC.