Podcasting is booming, but producing broadcast‑quality audio is a complex process. Recording raw audio often includes background noise, inconsistent loudness, and the wrong format. Professional podcasts need:
- Consistent loudness — EBU R128 loudness normalization (target -16 LUFS for podcasts)
- Noise reduction — remove background hum, hiss, and room tone
- Format conversion — WAV/FLAC to MP3 for distribution
- Metadata — title, artist, album, artwork
- Waveform visualization — for podcast players and social media
This guide shows you how to build a fully automated podcast audio processing pipeline using a declarative YAML template and a transpiler that generates the complete SQL migration. The pipeline turns a raw recording into a broadcast‑ready podcast episode — all driven by PostgreSQL triggers and pgmq.
Key takeaways
- One YAML file — define your entire pipeline in a single, version‑controlled file.
- Automatic SQL generation — the transpiler produces the exact PostgreSQL migration.
- Sequential processing — each step runs in order, with outputs passed to the next.
- Professional audio processing — loudness normalization, noise reduction, format conversion, metadata, and waveforms.
- Video support — extract audio from video files automatically.
- Visual graph — understand your pipeline at a glance with an SVG diagram.
- Per‑run grouping – all outputs for a single upload are stored under a unique
runIdfolder.
The Gap: From Raw Audio to Podcast-Ready
Most podcasters record audio in high‑quality formats (WAV, FLAC) and then manually process it: normalize loudness, remove noise, convert to MP3, add metadata, and create a waveform image. This is time‑consuming and inconsistent.
What if the pipeline could be fully automated — triggered by the upload itself, processing in the background, and delivering a complete podcast episode ready for distribution? And what if you could define that pipeline in a declarative YAML file that you can version, share, and reuse?
This guide shows you exactly how to build that pipeline.
Architecture Overview
Processed Audio → Public Folder → pg_notify → User Notified
The pipeline consists of:
- Supabase Storage — dedicated buckets for each processing stage:
audio-uploads,audio-temp-1,audio-temp-2, andaudio-processed(final). - PostgreSQL triggers — one trigger per step. Each fires on
INSERTintostorage.objectswhen a file appears in a specific bucket. - pgmq — message queue for job processing (uses the existing
renderqueue). - ffmpeglab-runner — executes FFmpeg commands for audio processing.
- pg_notify — real‑time status updates.
- YAML transpiler — reads the pipeline definition and generates the SQL migration + SVG graph.
Important: This pipeline uses the existing render and logpiece tables from the FFmpegLab server. It does not create new tables — it only adds the pipeline components.
What the Pipeline Delivers
| Output | Format | Location |
|---|---|---|
| Podcast Audio | MP3 (192kbps) with ID3v2 metadata | audio-processed/{userId}/{pipelineId}/{runId}/podcast/ |
| Waveform Image | PNG (1200×200) | audio-processed/{userId}/{pipelineId}/{runId}/waveforms/ |
| Metadata | Title, Artist, Album, Genre, Cover Art | ID3v2 tags (embedded in MP3) |
| Real‑time notifications | pg_notify channels | N/A |
| Job tracking | render table | Existing FFmpegLab table |
| Logs | logpiece table | Existing FFmpegLab table |
Prerequisites
- A Supabase project (cloud or self‑hosted).
- ffmpeglab-server and ffmpeglab-runner deployed (see setup guide).
- The
renderandlogpiecetables must already exist (created by the FFmpegLab server migrations). - Access to your Supabase database (psql or the Supabase SQL Editor).
- Deno installed to run the transpiler.
The YAML‑Driven Approach
While you can write the SQL directly, the recommended way is to use the YAML transpiler. This gives you:
- Declarative pipeline definition – define steps, triggers, and buckets in clean YAML.
- Automatic SQL generation – the transpiler produces the exact PostgreSQL migration.
- Visual pipeline graph – generate an SVG diagram of your pipeline with
--svg. - Reusable templates – share and version your pipeline definitions.
The transpiler is a single TypeScript file that reads your YAML and generates the SQL migration. It runs with Deno and has zero external dependencies (except yaml for parsing).
The YAML Template
Create a file called audio-pipeline.yaml with the following content. It defines the buckets, RLS policies, and each processing step in sequence. The runId section configures how the per‑run ID is generated — in this case, deterministically from the input file name.
name: "Audio Processing Pipeline" pipelineId: "audio-pipeline" runId: mode: "deterministic" template: "{baseFilename}" description: "Sequential audio processing: extract → normalize → waveform" version: "1.0.0" editor: compressionLevel: 23 preset: "medium" aspectRatio: "16:9" framerate: 30 opacity: 1.0 storage: output_bucket: "audio-processed" buckets: - name: "audio-uploads" public: false allowed_mime_types: - "audio/mpeg" - "audio/wav" - "audio/flac" - "video/mp4" - name: "audio-temp-1" public: false allowed_mime_types: - "audio/wav" - name: "audio-processed" public: true allowed_mime_types: - "audio/mpeg" - "image/png" rls_policies: - name: "Users can upload to their own audio folders" operation: "INSERT" role: "authenticated" condition: | (bucket_id IN ('audio-uploads', 'audio-temp-1', 'audio-processed')) AND (storage.foldername(name))[1] = auth.uid()::text steps: - id: "extract_audio" trigger: name: "handle_extract_audio" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'audio-uploads' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -ac 1 -ar 16000 -vn -f wav -y $OUTPUT_PATH inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/temp/{{baseFilename}}.wav" editor: output: "wav" preset: "fast" selectedCode: "custom" width: 0 height: 0 length: 0 compressionLevel: 0 next_bucket: "audio-temp-1" keep: false - id: "normalize_loudness" trigger: name: "handle_normalize_audio" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'audio-temp-1' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -af loudnorm=I=-16:LRA=11:TP=-1.5 -c:a libmp3lame -b:a 192k -f mp3 -y $OUTPUT_PATH inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/podcast/{{baseFilename}}.mp3" editor: output: "mp3" preset: "medium" selectedCode: "custom" compressionLevel: 23 next_bucket: "audio-processed" keep: true - id: "generate_waveform" trigger: name: "handle_waveform" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'audio-processed' AND NEW.name LIKE '%.mp3' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -filter_complex showwavespic=s=1200x200:colors=#FC6D26 -frames:v 1 -f image2 -y $OUTPUT_PATH inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/waveforms/{{baseFilename}}.png" editor: output: "png" preset: "medium" selectedCode: "custom" width: 1200 height: 200 next_bucket: "audio-processed" keep: true render: project_name: "audio-processing" status: "queued" public: false
The keep: true flag on the last step tells the transpiler to send the output directly to the final bucket (audio-processed). Intermediate steps use next_bucket to pass the result to the next step's trigger. The runId is computed deterministically from the input file name (using mode: "deterministic" and template: "{baseFilename}"). This ensures all steps in the sequential pipeline compute the same run ID, grouping all outputs for a single upload under one folder.
Running the Transpiler
Download the transpiler and the SVG generator:
curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/transpiler.ts
curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/svg.ts
Run the transpiler to generate the migration files:
deno run --allow-read --allow-write transpiler.ts audio-pipeline.yaml ./supabase/migrations
Add the --svg flag to also generate a visual graph of your pipeline:
deno run --allow-read --allow-write transpiler.ts audio-pipeline.yaml ./supabase/migrations --svg
The output will be:
UP: ./supabase/migrations/20260807120000_audio-pipeline.sql
DOWN: ./supabase/migrations/20260807120000_audio-pipeline_down.sql
SVG: ./supabase/migrations/20260807120000_audio-pipeline.svg
Apply the migration to your Supabase database:
psql -U postgres -d your_database -f ./supabase/migrations/20260807120000_audio-pipeline.sql
Visualising the Pipeline
The generated SVG gives you a clear overview of your pipeline. Steps marked with KEEP are green – their outputs are permanently stored in the final bucket. Edges are labelled with the bucket they use for data flow.
In the graph above, the steps run sequentially. The first step extracts audio and passes it to the second, which normalizes and passes to the third. The final step produces a waveform image and both the MP3 and PNG are stored in audio-processed. All steps share the same runId, so all outputs are grouped under {userId}/audio-pipeline/{runId}/.
Processing Logic Explained
When a file is uploaded to audio-uploads/{userId}/, the pipeline:
- Step 1 (extract_audio) — converts the file to WAV and uploads it to
audio-temp-1. - Step 2 (normalize_loudness) — applies loudness normalization and noise reduction, converts to MP3 with metadata, and uploads to
audio-processed(final bucket). - Step 3 (generate_waveform) — triggers on the MP3 uploaded to
audio-processed, generates a waveform image, and uploads it to the same bucket.
Each step's trigger condition matches the bucket that receives the previous step's output. The keep flag on the last two steps ensures the artifacts are permanently stored. Since runId is computed deterministically from the input file name, all steps get the same runId, grouping everything together.
Exact FFmpeg Commands
The YAML steps define the following FFmpeg commands using placeholders:
$MEDIA_1— The path to the downloaded input file (resolved by the runner).$OUTPUT_PATH— The temporary path for the output file (resolved by the runner).
1. Extract Audio (Step 1)
-ac 1— Convert to mono-ar 16000— Resample to 16kHz (speech-optimised)-vn— Drop any video stream-f wav— Output WAV format
2. Normalize Loudness (Step 2)
loudnorm=I=-16:LRA=11:TP=-1.5— EBU R128 loudness normalization target -16 LUFS (podcast standard)-c:a libmp3lame -b:a 192k— MP3 encoding at 192kbps-f mp3— Output MP3 format
3. Generate Waveform (Step 3)
showwavespic— Generates a waveform images=1200x200— Image size (1200×200 pixels)colors=#FC6D26— Waveform colour (FFmpegLab orange)-frames:v 1— Output a single frame (image)-f image2— Image output format
Configure ffmpeglab-runner
The runner needs to be configured to poll the render queue and execute the provided FFmpeg commands. The transpiler uses the existing render queue.
.env file or Docker Compose configuration.RENDER_QUEUE_NAME=render
apt-get install -y ffmpeg
# Alpine
apk add ffmpeg
Monitor the Pipeline
You can monitor the pipeline using SQL queries and notifications.
WHERE status = 'queued'
ORDER BY created_at DESC;
render table for job status.FROM "render"
WHERE project = 'audio-pipeline'
ORDER BY created_at DESC;
LISTEN render_status_channel;
LISTEN log_channel;
WHERE bucket_id = 'audio-processed'
ORDER BY created_at DESC;
Customising the Pipeline
Add Metadata to the MP3
Modify the command in the normalize_loudness step to include metadata tags:
Adjust Loudness Target
Change the I parameter in loudnorm:
- -18 — for music (more dynamic)
- -16 — for podcasts (recommended)
- -14 — for loud podcasts
Add Noise Reduction
Add the afftdn filter to the command:
Change Waveform Size or Color
Frequently Asked Questions (FAQ)
What audio processing does the pipeline perform?
The pipeline automatically normalizes loudness (EBU R128 with -16 LUFS target), reduces background noise using FFT-based denoising, converts formats, adds MP3 metadata, and generates waveform images for podcasts.
What FFmpeg filters are used?
The pipeline uses loudnorm for loudness normalization, afftdn for noise reduction, aformat for format conversion, ametadata for MP3 tags, and showwavespic for waveform visualization.
Can this handle videos as input?
Yes. The pipeline detects video files and extracts the audio track before processing. This makes it perfect for podcasters who record video interviews or screen recordings.
What podcast metadata is supported?
The pipeline supports title, artist (podcast name), album, genre, comment, and cover art (podcast artwork) in ID3v2 tags for MP3 files. You can customize these by modifying the metadata fields in the FFmpeg command.
Does this pipeline create new tables?
No. The pipeline uses the existing render and logpiece tables from the FFmpegLab server. It only adds storage buckets, RLS policies, the pgmq queue, and the trigger function — no table conflicts.
Final Word
You now have a fully automated podcast audio processing pipeline defined in YAML and generated via a transpiler. With PostgreSQL triggers, pgmq, and Supabase Storage, you get:
- Professional loudness normalization — EBU R128 compliant
- Noise reduction — clean, clear audio
- MP3 conversion — with ID3v2 metadata
- Waveform visualization — for podcast players
- Video support — extract audio from videos
- Real‑time notifications — know when processing is complete
- Full observability — monitoring views and logs
- Declarative YAML — version‑controlled, reusable pipeline definitions
- Per‑run grouping via deterministic
runId– all outputs for one upload stay together.
The pipeline is production‑ready, scalable, and configurable. It uses the existing render and logpiece tables from the FFmpegLab server, so there are no table conflicts — just pure, automated audio processing.