Home / Case Studies / Nomadic Wealth Academy
Case Study
AI Video & Carousel Content Generation Pipeline
One webhook in. Fully edited, captioned, avatar-narrated videos and matching carousels out — no editor, no studio, no manual assembly.
Client: The Nomadic Wealth Academy. Designed and built by Auto Core System.
At a glance
The system in four numbers.
Zero
Editors, voice actors or designers needed
1 call
Full content batch per webhook call
Word-level
Caption timing, not approximated
Video + carousel
Finished formats from one brief
Client: The Nomadic Wealth Academy, a content brand built by Brandon C.W. Johnson around remote income and travel. Auto Core System designed and built this pipeline.
Why automation wins here
Producing short-form video normally means a voice actor, an editor, a designer and a presenter. This pipeline replaces all four.
- Batch, not one-by-one. A human editor works one video at a time; this pipeline takes an entire batch of briefs through voice, video, captions and assembly in a single run — the workload doesn’t scale with headcount.
- Precision a human wouldn’t bother to match by hand. Manual caption sync is eyeballed against the audio; here captions are built directly from the voiceover’s word-level timestamp data, so timing is exact rather than approximate.
- No filming, ever. A consistent on-camera presenter normally means booking and filming a real person for every piece; the avatar and lip-sync layer produces that same presenter-led feel without a camera ever rolling.
- One brief becomes two deliverables. A manual process would produce a video and stop there; this pipeline generates a matching carousel from the same brief in the same run.
How it works
One webhook in, finished content out.
A single webhook call submits a full batch of content briefs, sourced directly from the client’s content workflow.
Each script gets a natural-sounding ElevenLabs voiceover, with word-level timestamp data returned alongside the audio.
That timing data is parsed into accurately synced caption lines, grouped by word count and duration, and a matching SRT file is generated.
Each piece of content is randomly assigned one of several AI presenter avatars, so output doesn’t feel repetitive.
Five distinct AI video-generation prompts are built per script — one per scene, each paired with its on-screen text — submitted to a model on Replicate and polled until every clip finishes rendering.
A separate avatar-specific voiceover segment is generated and run through a lip-sync model, so the assigned presenter appears to speak the hook naturally.
All clips, the avatar lip-sync clip and the voiceover are ingested into Shotstack and combined into a first-pass render with timing from the voiceover.
A final pass layers in the avatar lip-sync clip, the synced captions and all scene clips into one finished video.
In parallel, hook images are generated with Flux Schnell and assembled into a carousel-ready image set from the same content batch.
Finished video and carousel assets are sent to Telegram as soon as rendering completes, with backups pushed to Google Drive and Cloudinary.
Results & impact
What changed.
| Metric | Before | After |
| Script to finished video | Manual, multi-tool process | Fully automated, single pipeline |
| Caption accuracy | Manual or rough auto-sync | Word-level timed from voiceover data |
| Presenter / avatar content | Requires filming | AI avatar + lip-sync, generated automatically |
| Output formats | Video only | Video + carousel images from one input |
| Delivery | Manual export / sharing | Automatic delivery to Telegram + cloud backup |
The pipeline is in active production for The Nomadic Wealth Academy.
Tech stack
What it’s built on.
| Orchestration | n8n |
| Voice & timestamps | ElevenLabs |
| Video generation & lip-sync | Replicate (AI video + lip-sync models) |
| Video assembly & rendering | Shotstack |
| Image generation (carousels) | Flux Schnell |
| Asset backup & storage | Cloudinary, Google Drive |
| Delivery | Telegram API |
Content source and trigger: thenomadicwealthacademy.com