Video understanding is one of the most powerful applications of AI. Being able to automatically detect objects, faces, and scenes in videos opens up a world of possibilities — from content moderation and search to analytics and accessibility.
FFmpeg's DNN filters — dnn_detect and dnn_classify — make this possible. With pre‑trained models like YOLO, ResNet, and face detection models, you can label every frame of your video with rich metadata.
This guide shows you how to build a fully automated video labeling pipeline using a declarative YAML template and a transpiler that generates the complete SQL migration. The pipeline:
- Detects objects and faces in uploaded videos
- Classifies scenes and emotions
- Generates JSON metadata with bounding boxes and labels
- Stores results alongside the original video
Key takeaways
- One YAML file — define your entire pipeline in a single, version‑controlled file.
- Automatic SQL generation — the transpiler produces the exact PostgreSQL migration.
dnn_detect— object and face detection with bounding boxesdnn_classify— classification of scenes, emotions, and more- YOLO, ResNet, and face detection — proven models for video understanding
- JSON metadata — results stored for search, analytics, and further processing
- Visual graph — understand your pipeline at a glance with an SVG diagram
- Per‑run grouping – all outputs for a single upload are stored under a unique
runIdfolder.
The Gap: Video Understanding at Scale
Video is the most data‑rich medium we have. But without understanding what's in the video, it's just pixels. Manually labeling videos is impossible at scale. Automated video labeling using DNN filters makes it practical:
- Content moderation — detect inappropriate content automatically
- Search and discovery — find videos by content (e.g., "videos with dogs")
- Analytics — understand what's in your video library
- Accessibility — auto‑generate alt text and descriptions
What if the pipeline could be fully automated — triggered by the upload itself, running DNN detection and classification in the background? And what if you could define that pipeline in a declarative YAML file that you can version, share, and reuse?
This guide shows you exactly how to build that pipeline.
Architecture Overview
dnn_detect → Bounding Boxes → dnn_classify → Labels
JSON Metadata → Public Folder → pg_notify → User Notified
The pipeline consists of:
- Supabase Storage — two buckets:
dnn-uploads(per‑user) andpublic-processed(per‑user). - RLS policies — restrict access to each user's own folders.
- PostgreSQL trigger — fires on
INSERTintostorage.objects. - pgmq — message queue for job processing (uses the existing
renderqueue). - ffmpeglab-runner — runs DNN detection and classification.
- pg_notify — real‑time status updates.
- YAML transpiler — reads the pipeline definition and generates the SQL migration + SVG graph.
Important: This pipeline uses the existing render and logpiece tables from the FFmpegLab server. It does not create new tables — it only adds the pipeline components.
What the Pipeline Delivers
| Output | Format | Location |
|---|---|---|
| Detection Metadata | JSON (bounding boxes, labels, confidence) | public-processed/{userId}/{pipelineId}/{runId}/labels/ |
| Labeled Video | MP4 (with bounding boxes overlaid) | public-processed/{userId}/{pipelineId}/{runId}/labeled/ |
| Real‑time notifications | pg_notify channels | N/A |
| Job tracking | render table | Existing FFmpegLab table |
| Logs | logpiece table | Existing FFmpegLab table |
Prerequisites
- A Supabase project (cloud or self‑hosted).
- ffmpeglab-server and ffmpeglab-runner deployed.
- FFmpeg compiled with DNN support (OpenVINO or TensorFlow).
- DNN model files downloaded (see Model Management).
- The
renderandlogpiecetables must already exist (created by the FFmpegLab server migrations). - Access to your Supabase database (psql or the Supabase SQL Editor).
- Deno installed to run the transpiler.
The YAML‑Driven Approach
While you can write the SQL directly, the recommended way is to use the YAML transpiler. This gives you:
- Declarative pipeline definition — define steps, triggers, and buckets in clean YAML.
- Automatic SQL generation — the transpiler produces the exact PostgreSQL migration.
- Visual pipeline graph — generate an SVG diagram of your pipeline with
--svg. - Reusable templates — share and version your pipeline definitions.
The transpiler is a single TypeScript file that reads your YAML and generates the SQL migration. It runs with Deno and has zero external dependencies (except yaml for parsing).
The YAML Template
Create a file called labeling-pipeline.yaml with the following content. It defines the buckets, RLS policies, and each processing step. The runId section configures how the per‑run ID is generated — in this case, deterministically from the input file name.
name: "Video Labeling with DNN Filters" pipelineId: "video-labeling" runId: mode: "deterministic" template: "{baseFilename}" description: "Detect objects, faces, and scenes using FFmpeg DNN filters" version: "1.0.0" editor: compressionLevel: 23 preset: "medium" aspectRatio: "16:9" framerate: 30 opacity: 1.0 output: "mp4" storage: output_bucket: "public-processed" buckets: - name: "dnn-uploads" public: false allowed_mime_types: - "video/mp4" - "video/quicktime" - "video/x-msvideo" - "video/webm" - "video/mpeg" - name: "public-processed" public: true allowed_mime_types: - "video/mp4" - "application/json" rls_policies: - name: "Users can upload to their own folder" operation: "INSERT" role: "authenticated" condition: | bucket_id = 'dnn-uploads' AND (storage.foldername(name))[1] = auth.uid()::text - name: "Users can download from their own folder" operation: "SELECT" role: "authenticated" condition: | bucket_id = 'dnn-uploads' AND (storage.foldername(name))[1] = auth.uid()::text - name: "Public read access to processed media" operation: "SELECT" role: "anon" condition: | bucket_id = 'public-processed' - name: "Service role can manage processed media" operation: "ALL" role: "service_role" condition: | bucket_id = 'public-processed' - name: "Users can read their own processed media" operation: "SELECT" role: "authenticated" condition: | bucket_id = 'public-processed' AND (storage.foldername(name))[1] = auth.uid()::text steps: # Step 1: Object/Face Detection (dnn_detect) - id: "detect_objects" trigger: name: "handle_detect" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'dnn-uploads' AND NEW.metadata->>'mimetype' LIKE 'video/%' command: -i $MEDIA_1 -vf "dnn_detect=dnn_backend=$DNN_BACKEND:model=$DETECT_MODEL:input=data:output=detection_out:confidence=$DETECT_CONFIDENCE:labels=$DETECT_LABELS,showinfo" -f null - inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labels/{{baseFilename}}_detections.json" editor: output: "json" preset: "medium" selectedCode: "custom" next_bucket: "dnn-uploads" keep: false # Step 2: Scene/Emotion Classification (dnn_classify) - id: "classify_scenes" trigger: name: "handle_classify" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'dnn-uploads' AND NEW.name LIKE '%.json' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -vf "dnn_classify=dnn_backend=$DNN_BACKEND:model=$CLASSIFY_MODEL:input=data:output=$CLASSIFY_OUTPUT:confidence=$CLASSIFY_CONFIDENCE:labels=$CLASSIFY_LABELS,showinfo" -f null - inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labels/{{baseFilename}}_classifications.json" editor: output: "json" preset: "medium" selectedCode: "custom" next_bucket: "dnn-uploads" keep: false # Step 3: Overlay bounding boxes on video - id: "overlay_labels" trigger: name: "handle_overlay" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'dnn-uploads' AND NEW.name LIKE '%.json' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -vf "dnn_detect=dnn_backend=$DNN_BACKEND:model=$DETECT_MODEL:input=data:output=detection_out:confidence=$DETECT_CONFIDENCE:labels=$DETECT_LABELS,drawbox=x=1005:y=813:w=81:h=92:color=red" -c:v libx264 -crf 18 -y $OUTPUT_PATH inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labeled/{{baseFilename}}_labeled.mp4" editor: output: "mp4" preset: "medium" selectedCode: "custom" next_bucket: "public-processed" keep: true render: project_name: "video-labeling" status: "queued" public: false
The keep: true flag on the last step tells the transpiler to send the output directly to the final bucket (public-processed). Intermediate steps use next_bucket to pass the result to the next step's trigger. The runId is computed deterministically from the input file name (using mode: "deterministic" and template: "{baseFilename}"). This ensures all steps in the sequential pipeline compute the same run ID, grouping all outputs for a single upload under one folder.
Running the Transpiler
Download the transpiler and the SVG generator:
Run the transpiler to generate the migration files:
Add the --svg flag to also generate a visual graph of your pipeline:
The output will be:
Apply the migration to your Supabase database:
Visualising the Pipeline
The generated SVG gives you a clear overview of your pipeline. Steps marked with KEEP are green – their outputs are permanently stored in the final bucket. Edges are labelled with the bucket they use for data flow.
In the graph above, the steps run sequentially. The first step detects objects and faces, the second classifies scenes and emotions, and the final step overlays bounding boxes on the video. All steps share the same runId, so all outputs are grouped under {userId}/video-labeling/{runId}/.
FFmpeg Commands
The YAML steps define the following FFmpeg commands using placeholders:
$MEDIA_1— The path to the downloaded input file (resolved by the runner).$OUTPUT_PATH— The temporary path for the output file (resolved by the runner).$DNN_BACKEND,$DETECT_MODEL,$CLASSIFY_MODEL— Environment variables for model configuration.
1. Object / Face Detection (dnn_detect)
dnn_detect— FFmpeg's object detection filterdnn_backend=openvino— DNN backend (openvino, tensorflow, native)model— Path to the model fileinput=data— Input tensor nameoutput=detection_out— Output tensor nameconfidence=0.6— Confidence thresholdlabels— Path to labels fileshowinfo— Prints detection results to console
2. Classification (dnn_classify)
dnn_classify— FFmpeg's classification filteroutput=prob_emotion— Output tensor for probabilitiesconfidence=0.3— Confidence threshold
3. Overlay Detection Results (drawbox)
drawbox— FFmpeg filter for drawing rectanglesx, y, w, h— Position and size of the bounding boxcolor=red— Color of the bounding box-c:v libx264 -crf 18— High-quality encoding-y— Overwrite output file
Model Management
Directory Structure
Downloading Models
Configure ffmpeglab-runner
The runner needs to be configured to poll the render queue and execute the provided commands. The transpiler uses the existing render queue.
.env file or Docker Compose configuration.# 2. For each job:
# a. Download the input file
# b. Parse the 'commands' array
# c. Execute each command:
# - Run dnn_detect (extract JSON metadata)
# - Run dnn_classify (extract JSON metadata)
# - Run overlay (generate labeled video)
# d. Upload outputs to public-processed
# e. Update render table
# f. Mark job complete and delete from queue
Monitor the Pipeline
You can monitor the pipeline using SQL queries and notifications.
render table for job status.Customising the Pipeline
Use a Different Detection Model (YOLO)
Replace DETECT_MODEL and DETECT_LABELS with YOLO paths:
Update the YAML command:
Customise Overlay Styling
Modify the drawbox parameters in the overlay step:
Add Multi‑Frame Detection
To detect objects across multiple frames, use the dnn_detect filter with frame caching:
Frequently Asked Questions (FAQ)
What does the video labeling pipeline do?
The pipeline automatically detects and labels objects, faces, and scenes in uploaded videos using FFmpeg's dnn_detect and dnn_classify filters. It generates bounding boxes, labels, and confidence scores, and stores them as JSON metadata alongside the video.
What models are used?
The pipeline supports multiple models including YOLO for object detection, ResNet for classification, and face detection models. The default configuration uses OpenVINO models for face detection (face-detection-adas-0001), classification (emotions-recognition-retail-0003), and can be extended for general object detection.
What is the output format?
The pipeline outputs a JSON file containing all detection results: frame number, bounding box coordinates, label/class name, confidence score, and timestamp. This can be used for search, analytics, or further processing.
Can I use custom models?
Yes. You can use any OpenVINO or TensorFlow model that works with FFmpeg's dnn_detect or dnn_classify filters. You'll need to provide the model files and configure the input/output tensor names.
Does this pipeline create new tables?
No. The pipeline uses the existing render and logpiece tables from the FFmpegLab server. It only adds storage buckets, RLS policies, and the trigger function — no table conflicts.
Final Word
You now have a fully automated video labeling pipeline defined in YAML and generated via a transpiler. With PostgreSQL triggers, pgmq, and Supabase Storage, you get:
- Object and face detection — bounding boxes with confidence scores
- Scene classification — labels for what's in the video
- JSON metadata — rich, structured data for search and analytics
- Labeled video output — visual overlay of detection results
- Real‑time notifications — know when processing is complete
- Full observability — monitoring views and logs
- Declarative YAML — version-controlled, reusable pipeline definitions
- Per‑run grouping via deterministic
runId– all outputs for one upload stay together.
The pipeline is production‑ready, scalable, and extensible — you can swap in any DNN model for your specific use case.