Safe City vision-language system
A vision-language model for city safety monitoring: describing what is happening in a camera feed, not only detecting objects inside it.
The challenge
Classic detection tells an operator that a box appeared on screen. It cannot tell them what the box is doing, and a room of operators cannot watch every feed. Safe-city work needs a system that understands a scene well enough to describe an event worth a human's attention.
Our approach
Understand scenes, not objects
A vision-language approach so the system can produce a description of an event, not only a class label.
Run where the cameras are
Inference designed for on-premise hardware, so video does not have to leave the operator's network.
Rank for the operator
The output is a shortlist of events worth looking at, with the reason attached.
Screenshots of the deployed interface are not published: the deployment is covered by a partner NDA.