Home / Case Studies / Nomadic Wealth Academy
Case Study

AI Video & Carousel Content Generation Pipeline

One webhook in. Fully edited, captioned, avatar-narrated videos and matching carousels out — no editor, no studio, no manual assembly.

Client: The Nomadic Wealth Academy. Designed and built by Auto Core System.

At a glance

The system in four numbers.

Zero
Editors, voice actors or designers needed
1 call
Full content batch per webhook call
Word-level
Caption timing, not approximated
Video + carousel
Finished formats from one brief

Client: The Nomadic Wealth Academy, a content brand built by Brandon C.W. Johnson around remote income and travel. Auto Core System designed and built this pipeline.

Why automation wins here

Producing short-form video normally means a voice actor, an editor, a designer and a presenter. This pipeline replaces all four.

  • Batch, not one-by-one. A human editor works one video at a time; this pipeline takes an entire batch of briefs through voice, video, captions and assembly in a single run — the workload doesn’t scale with headcount.
  • Precision a human wouldn’t bother to match by hand. Manual caption sync is eyeballed against the audio; here captions are built directly from the voiceover’s word-level timestamp data, so timing is exact rather than approximate.
  • No filming, ever. A consistent on-camera presenter normally means booking and filming a real person for every piece; the avatar and lip-sync layer produces that same presenter-led feel without a camera ever rolling.
  • One brief becomes two deliverables. A manual process would produce a video and stop there; this pipeline generates a matching carousel from the same brief in the same run.
How it works

One webhook in, finished content out.

A single webhook call submits a full batch of content briefs, sourced directly from the client’s content workflow.
Each script gets a natural-sounding ElevenLabs voiceover, with word-level timestamp data returned alongside the audio.
That timing data is parsed into accurately synced caption lines, grouped by word count and duration, and a matching SRT file is generated.
Each piece of content is randomly assigned one of several AI presenter avatars, so output doesn’t feel repetitive.
Five distinct AI video-generation prompts are built per script — one per scene, each paired with its on-screen text — submitted to a model on Replicate and polled until every clip finishes rendering.
A separate avatar-specific voiceover segment is generated and run through a lip-sync model, so the assigned presenter appears to speak the hook naturally.
All clips, the avatar lip-sync clip and the voiceover are ingested into Shotstack and combined into a first-pass render with timing from the voiceover.
A final pass layers in the avatar lip-sync clip, the synced captions and all scene clips into one finished video.
In parallel, hook images are generated with Flux Schnell and assembled into a carousel-ready image set from the same content batch.
Finished video and carousel assets are sent to Telegram as soon as rendering completes, with backups pushed to Google Drive and Cloudinary.
Results & impact

What changed.

The pipeline is in active production for The Nomadic Wealth Academy.

Tech stack

What it’s built on.

Content source and trigger: thenomadicwealthacademy.com

Want something like this built for you?

Book a free 30-minute call and we’ll map what it would take for your business.