Building an AI Surveillance Video Analyzer with Codex to Solve a Burglary

The author describes building a surveillance video analysis system using Codex and GPT-luna-5.6 to process 20+ hours of footage, extracting frames, filtering static frames, creating 16-frame grids, and using natural language prompts to identify suspicious activity, with resume capability.

IT Services Circle
IT Services Circle
IT Services Circle
Building an AI Surveillance Video Analyzer with Codex to Solve a Burglary

Problem: 20+ Hours of Surveillance Footage After a Burglary

A friend's house was burglarized. The surveillance system recorded over 20 hours of footage across 400+ video files, but the exact time of the theft was unknown. Manual review at high speed was exhausting and required constant attention to spot people and any suspicious objects they carried.

Solution Overview: Automated Pipeline with AI Analysis

The author built a Surveillance Video AI Analyzer using Codex (an AI coding assistant) to automate the review. The pipeline consists of four main stages:

Frame extraction: Sample frames at a fixed interval (e.g., 1 frame per second) from each video file recursively.

Static frame filtering: Remove frames where the scene is unchanged (static background) to reduce the number of frames sent to the AI. This is done by comparing consecutive frames and discarding those below a difference threshold.

Grid composition: Group the remaining frames into 16-frame grids (4x4) to maximize the number of frames per API request.

AI analysis: Send each grid (4 grids per request, up to 4 concurrent requests = 16 frames in parallel) to the GPT-luna-5.6 model, chosen for its low cost.

Flowchart of the video analysis pipeline
Flowchart of the video analysis pipeline

AI Analysis Details

The tool provides an "Analysis Intent" input box where the user describes the target in natural language. For the burglary case, the prompt was:

"Find shots where a person wearing a backpack appears."

The model returns for each grid:

Whether it matches the intent (hit/miss)

The reasoning for the decision

The specific sub-grid indices (1-16) that triggered the hit

Matched frames are highlighted and collected in an "AI Hits" page. Clicking a grid shows the enlarged view with the relevant sub-frames marked and the model's explanation. The system also records the source video file and the time range each grid covers, enabling quick navigation back to the original footage. Users can bookmark important frames for later review.

16-frame grid example from author's X post
16-frame grid example from author's X post

Resumable Long-Running Tasks

Because the analysis runs over many videos and network or API interruptions can occur, the tool saves its progress (task state) automatically. If the process stops, it can resume from the last saved point without re-processing already analyzed videos.

Comparison with Alternative Approaches

The author acknowledges that this solution is not necessarily optimal. Traditional computer vision methods like YOLO object detection are purpose-built for detecting people and objects in video and are widely used in commercial surveillance products. The LLM-based approach trades off latency and cost for flexibility in natural-language query definition and zero training data requirement.

Key Takeaways

LLMs can be combined with simple frame sampling and differencing to create a functional video search tool quickly.

Grid packing and concurrent requests reduce API calls and cost.

State persistence enables reliable processing of large datasets.

The project demonstrates applying AI to a concrete real-world problem rather than just chat or code generation.

The author is considering open-sourcing the tool if there is sufficient interest.

AI hit highlights interface
AI hit highlights interface

Code example

来源丨
经授权转
自
轩辕的编程宇宙(ID:xuanyuancoding)
作者丨
轩辕之风
大家好,我是轩辕。
最近发生了一件事。
朋友家里被盗了,不过好在有监控。
因为不能确定准确的时间,所以需要排查的时间段累计20多个小时,总共400多个视频。
于是他有点犯难了,20多个小时,一段一段看,不得累死。
有人说,开倍速啊,拖进度条啊。
朋友不是没试过,但看了十多分钟就累得够呛了,这玩意儿需要精神高度集中,不仅要看人,有人出现后还要停下来看看有没有携带可疑的东西。
于是朋友找到我:你不是在折腾AI吗,能不能让AI给我自动化分析一下啊。
于是我想了想,用Codex手搓了一个
监控视频AI分析器
出来。
让程序去处理这些视频,再让AI帮忙筛选画面。没跑多久,就定位到了关键画面。
AI这回是帮了大忙了。
接下来说说我是怎么做的。
1. 整体过程
整个处理过程可以看下面这张图。
指定要分析的视频文件夹,系统自动递归遍历所有视频文件,依次处理。
第一步:根据设定是时间间隔从视频抽帧,比如1秒一帧。
第二步:过滤。监控视频大量时间画面都是静态不动的,把这些静态不动的帧去掉,减少分析工作量。
第三步:把过滤后的图片拼成16宫格,交给AI分析。
我选择的是GPT-luna-5.6模型来进行分析,因为这玩意儿很便宜。
抽帧后拼成16宫格的实际效果,来自我的X原帖
程序一次请求发送4张宫格图,同时最多跑4路请求,相当于16帧并发分析。
2. AI分析
工具里有一个“分析意图”输入框,可以直接用自然语言描述你想找什么。
比如:“寻找里面有背背包的人出没的镜头。”
模型会逐张返回是否命中、判断原因,以及命中的宫格编号。
命中的图片会高亮,并集中放到“AI 命中”页面。打开大图,还能看它具体指的是哪几格、为什么选中。
工具会保留来源视频信息,在输出记录里保存每张宫格图的采样时间范围,方便回查。需要重点看的图片,也能收藏起来。
还有一个跑长任务时很实用的功能:
中断了,可以继续。
因为网络、AI中转站等原因导致分析中断的时候,会保存任务状态,然后可以继续,不用重头再跑。
3、最后
估计有一些人会说:你这方案不行,我知道另一个xx方案更好···
确实,这套方案解决了我眼前的问题,但它未必是这个需求的最佳方案。
像YOLO这样的目标检测技术,本来就能用来检测图片、视频中的目标。行业里也早已有成熟的检测方案和监控分析产品。
我分享这个,并非想说这是最好的解决方案,而是提供一种方法供大家参考了解。
觉得有用,收藏一下。
觉得没用,划走即可。
除此之外,我更想分享的是,我们可以用AI解决日常工作生活中问题的这种思想。
AI不只是写写文章,做一些网站系统,它是真的可以用到我们生活中来。
如果你做过类似的项目,有更省事、更稳定的方案,也欢迎在评论区交流,大家一起少走弯路。
如果你对这个监控视频AI分析器感兴趣,可以在评论区留言。感兴趣的朋友多的话,我可以考虑整理一下,把它开源出来。
好了,今天就分享到这里。
如果觉得还不错的话,求个点赞。
路过的朋友欢迎点个关注呗,及时获得下期文章推送~
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

prompt engineeringvideo analysisframe extractionCodexsurveillanceGPT-luna-5.6grid compositionresumable processingstatic frame filtering
IT Services Circle
Written by

IT Services Circle

Delivering cutting-edge internet insights and practical learning resources. We're a passionate and principled IT media platform.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.