From Video Analysis Pipeline to Production System: Architecture, Contracts, and Guardrails

This article details the system architecture for production-ready video AI analysis, covering task contracts with three distinct identifiers, media and model isolation, capacity estimation using Little's Law, retry and recovery rules with outbox pattern, evidence retention requirements, and a phased gray-release strategy for two scenarios: engineering quality inspection and intelligent security.

Chengwu Tech Stack
Chengwu Tech Stack
Chengwu Tech Stack
From Video Analysis Pipeline to Production System: Architecture, Contracts, and Guardrails

Overview

This fourth lesson in the AI video quality inspection series moves from an analysis pipeline to a deliverable system architecture. A single video may process successfully, but real projects must handle duplicate submissions, corrupted videos, rate limits, task cancellation, manual review, and delivery confirmation. The article emphasizes defining responsibility boundaries first, then deciding which modules run in-process and which need separation based on throughput, recovery needs, and team capacity.

Logical Architecture

The logical architecture decouples video processing, model requests, and result delivery. Media and evidence are not shuffled across layers, and model observations do not directly become irreversible business actions. Logical layers do not mandate six separate services; a minimum viable version can be a task API, a worker process, a persistent task store, and controlled file storage. Media processing and model execution are split only when independent scaling or fault isolation is required.

01 Task Contract Before Deployment

Required Inputs

Media reference (must pass authorization; arbitrary URLs are not trusted)

Analysis target

Request idempotency key

Business correlation ID

Model keys stay in server-side controlled configuration, never passed via frontend or model messages.

Three Identifiers — Do Not Mix

request_id : troubleshooting a single API request

task_id : represents the persistent analysis task

idempotency_key : identifies the same business submission; validated within tenant/caller scope and bound to a fingerprint of key inputs. Same key with different video or target must raise a conflict, not return an old report.

Task Lifecycle

Acceptance returns a task ID immediately; results arrive via query or callback. Acceptance only means the task is logged, not completed. Internal states: queued, running, retry_wait. Final states: completed, partial, failed, cancelled. Processing status and observation conclusions (e.g., person presence) are two separate dimensions.

Result Contract

Retains execution_status, person_presence, scope, evidence_refs, missing_ranges, error_code, and versions. When no usable observation exists, person_presence stays uncertain. Whether to continue, escalate to human, or generate a candidate event is decided by business rules separately. External responses are concise; internal evidence is complete. Vendor raw exceptions, full headers, or keys are never passed to callers.

02 Media, Model, and Capacity Controlled Separately

Media Side: Define Analysis Scope First

Lesson 1 processed only the first 15 seconds for teaching; full videos need a window list, then per-window sampling with coverage recording. Overlap depends on boundary-event miss risk; overlapping frames must keep deduplicable source IDs. File entry limits size, duration, resolution, and supported formats. Decoder processes have resource and timeout boundaries. Failed videos enter an explicit error state — zero frames are not treated as "no target." External media downloaders and decoders need independent network access limits and process isolation.

Model Side: Adaptation, Rate Limiting, Versioning in One Layer

The adaptation layer handles message format, image transfer, structured output, and platform parameter differences; it logs call ID, latency, usage, and error categories. Model upgrades or vendor switches require fixed-sample regression first — no silent fallback to unverified models. Rate limiting accounts for request count, concurrency, and image/token budget. The task queue handles waiting; workers must not increase concurrency unboundedly to clear backlogs. Temporary image URLs are short-lived; logs and tracing avoid storing long-accessible media links.

Capacity Estimation Example

Assumptions: peak arrival 0.5 tasks/sec, average task occupies a worker slot for 8 seconds → average ~4 concurrent slots (Little's Law: λ × W). At 70% target utilization → ~6 slots. This is a steady-state estimate; validation requires high-percentile latency, burst traffic, and platform quota stress tests. If each task averages 2 model calls, that load equals ~1 model request/sec. Multi-image scale, retries, and agent resampling change this — cost must be calculated from real usage and platform billing rules, not model self-reports or fixed prices. Sampling output ≠ valid observation; capacity and cost optimization must monitor missampling, misjudgment, and review rates, not just fewer images or faster returns.

03 Rules for Retry, Recovery, and Delivery

Task Retry ≠ Business Action Retry

A failed model batch can be retried within budget. If results are saved but notification not delivered, retry notification delivery — do not re-run the whole video. Saving results and delivery records in the same transaction, then sending via a delivery process, is a common outbox design . Receivers deduplicate by stable task and result version; senders track pending, delivered, and pending-verification. After a request timeout, check receipt or query per agreement — "no response" ≠ "not executed." Actions lacking idempotency or verification go to manual handling.

Recovery: Validate Execution Validity First

Executors claim tasks with an attempt ID or lease version. Before saving results, verify the task is not cancelled and not taken over by a new attempt, preventing an old worker from overwriting new results after a timeout. Cancellation must propagate to nodes and external call boundaries; calls that cannot be interrupted immediately must not submit stale results upon return. Recovery relies on persistent task records, accessible evidence, and fixed config versions. Checkpoints assist image-state recovery but do not replace business idempotency, external receipts, and valid execution rights checks. Re-analysis produces a new result version, never overwriting the old report's evidence chain.

Minimum Evidence Retention

Retain: original video ID/content digest, sampling strategy, frame numbers/timestamps, model and prompt versions, raw response references, validation results, aggregation strategy, human modifications, delivery receipts. Not all raw data forever — define clear retention periods, access permissions, and expiry cleanup. Raw model output is not the sole truth; human review changes must keep original conclusion, modification reason, operator, and timestamp. Business reads current valid version; audit can review history. Recoverable means knowing what work is confirmed done, what is pending, and who has the right to continue — not re-running from scratch hoping results coincidentally match.

04 Two Scenarios: Reuse Capabilities, Not Actions

Engineering Video Quality Inspection

Map each conclusion to a check item. Example: verify a job segment contains personnel operation footage. Person appearance is only one observation condition — it does not mean the operation was fully or correctly completed. Check items define required time range, image quality, acceptable evidence; insufficient material triggers supplement or review, not auto-pass with "person present." Tasks are usually async queued; focus on report completeness, evidence review, explainable failure reasons, and human handling loop. Do not declare the whole video analysis passed just because all interfaces returned success.

Intelligent Security

Clarify timeliness first, then place the large model. For fast entry-event detection, use evaluated basic detection + temporal rules to generate candidate events, then vision model adds scene description, assists denoising and explanation. Only when explicitly supporting the required timeliness and risk do model results enter formal trigger rules. "A frame has a person" cannot alone prove "person entered restricted zone"; entry events need zone definition, before/after temporal or trajectory evidence. Cloud model unavailable → explicitly report capability degradation and follow established guard strategy; never silently drop candidate events. Observation results, alarm candidates, human confirmation, and device actions are distinct stages. Especially with on-site devices, permissions, confirmation, and safety interlocks must be independent of model text.

05 Acceptance and Gray Release

Pre-Launch Acceptance

Functional : video input, frame source, per-frame judgment, aggregation, review, result query, receipt — full chain works; one report answers what was seen, what wasn't, and why.

Exception : network break, rate limit, corrupted video, output truncation, duplicate submission, task cancellation, worker restart all have defined outcomes. Especially test partial and failed; frontend must not uniformly display them as "no issue found."

Effect : group by day/night, near/far targets, occlusion, site differences; measure false positive, false negative, uncertain rates, human burden. Thresholds set by specific business consequences and sample evidence — not a single accuracy number.

Operational : stress test queue wait, end-to-end latency, model call volume, retry rate, per-task cost; verify backup recovery, expired evidence handling, notification deduplication, version rollback.

Safe Gray-Release Sequence

Shadow run : produce observations only, no formal actions, compare with human results.

Assisted review : users see evidence and suggestions; record adoption and corrections.

Limited auto-processing : only for verified, low-risk scenarios with clear rollback. Each step has entry criteria and stop conditions.

Rollback: simultaneously fix the verified combination of model, prompt, sampling, and aggregation strategy. New tasks switch to the old combination; in-progress tasks complete per their bound version or are explicitly cancelled. Rollback config does not undo external actions already taken; independent verification and compensation processes may be needed.

Closing Path

The four lessons form a gradual implementation path: understand the scene, define judgments, organize automated flows, connect execution with business responsibility. Autonomy can increase gradually, but evidence, boundaries, and failure handling must exist from day one.

References

LangGraph persistence design

Qwen vision input capabilities and limits

LangChain structured output mechanism

Verification date: 2026-09-03. This article is a design proposal; sample inputs, parameters, and results are for rule illustration only — no real model effect testing or production integration was performed.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

system architecturegray releasecapacity planningtask contractidempotency keyevidence retentionretry recoveryvideo AI analysis
Chengwu Tech Stack
Written by

Chengwu Tech Stack

A powerful mindset is a lifelong treasure!

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.