This post was inspired by and builds on an analysis originally published in Chinese by 差不多小姐 on the WeChat Official Account 小小创业经 (Sept 15, 2026). Link at the end.
There’s a video circulating right now that’s making the rounds in every “make money with AI” corner of the internet:
A stick figure. Simple images, one after another. No face, no camera, no animation skills.
7.5 million views.
And the stat that broke people’s brains: when the video was made, the channel was two months old, had published only 14 videos, and had 137,000 subscribers — 14.5 million total views.
The video’s title? “Claude Code + YouTube = $77,000/month.”
Cue a thousand people opening Claude and asking it to write YouTube scripts.
I’ve spent time pulling this workflow apart, and I want to give you the version the hype video won’t: what’s genuinely clever here, what the numbers actually imply, what the workflow leaves out — and the one production change that matters more than any tool you’ll pick.
First, Let’s Talk About That $77,000
It’s not backed by a dashboard screenshot, so treat it as marketing.
Here’s the math you can check. A faceless “explainer facts” channel pulling millions of views typically monetizes through:
- AdSense: faceless channels in business/tech niches often see RPMs of $2–8, sometimes higher in finance
- Affiliate links and sponsorships: often 2–5x AdSense for channels that know their audience
- Selling the method itself: which, notably, is exactly what this creator is doing by posting the “here’s my workflow” video
That last bullet isn’t a knock — it’s the most important clue. When someone’s video about making money is itself the monetization, adjust your priors accordingly. That doesn’t mean the channel is fake. It means the workflow is real, the specific numbers are unaudited, and the business model of the video is you.
Still here? Good. Because buried under the clickbait is a genuinely well-engineered production pipeline. Let’s get into it.
The One Change That Beats Every Tool: Audio First
Here’s how most people build an AI video, and it’s wrong:
- Take the script
- Generate an image for each sentence
- Generate the voiceover
- Try to align everything in an editor
Step 4 is where dreams die. The image cuts away mid-sentence. The voiceover pauses for two seconds while the screen sits frozen. Frames flash by in one second before anyone can read them. Individually the images look great; stitched together, the video feels off in a way viewers can’t articulate but absolutely feel.
The creator’s fix is deceptively simple: lock the audio first, then let visuals follow the sound.
Voiceover has natural pacing — pauses, emphasis, turns. Those pauses are your cut points. The audio literally tells you where scenes should change.
This is the most valuable idea in the entire workflow, and it costs $0. It’s also completely tool-agnostic: it works whether you generate images with Midjourney, Nano Banana, or a pencil.
The Full Pipeline, Step by Step
Step 1: Script — with a garbage filter
Paste a master prompt into Claude, pick one of 5 topic options, and get a complete script for a 5–15 minute video.
Now the part everyone skips: AI scripts fail in a specific way. They produce fluent, correct-sounding filler that viewers retain nothing from. Before voiceover, audit the script against three questions:
- Do the first 30 seconds actually hook someone?
- Is there a new piece of information or a story turn at regular intervals?
- If you deleted half of it, would the video lose anything?
If cutting a big chunk changes nothing, the script is water. Fix it before generating a single image — a weak script can’t be rescued by 1,000 pretty frames.
Step 2: Voiceover — before any images
Generate the full voiceover in ElevenLabs (the original creator used a voice named “Raunak”; any consistent voice works). Then finish it: fix mispronunciations, adjust pacing, rewrite awkward sentences.
Don’t proceed until the audio is clean. Every timestamp, image count, and edit beat downstream inherits from this file. Change the audio later and you may redo everything after it.
Step 3: Extract the timing map
Upload the audio to a transcription tool that produces timestamped transcripts. The original workflow used FaziScribe; alternatives include Whisper-based tools (subtitles via CapCut, Descript, or whisperx if you’re technical) — anything that gives you word- or sentence-level timestamps.
The output converts audio pauses into a concrete cut list: sentence one pauses at second 2, sentence two at second 4, the next at second 7. Those are your scene boundaries. This single step replaces hours of scrubbing through audio by ear.
Send the timestamped transcript back to Claude.
Step 4: One image prompt per scene
Claude converts each narration segment into an image prompt, aligned to its timestamp. A long video needs 100+ scenes; have it work in batches until every segment has one.
Here’s the trap the original video glosses over: the hard problem isn’t generating images, it’s consistency. A character’s hands, the direction faces point, background color temperature — if these drift across 100+ frames, viewers won’t say what’s wrong, but they’ll feel it and leave. Whatever tool you use, force it into a single visual system: reuse a detailed style descriptor in every prompt, fix a seed if the tool supports it, generate a reference character first and describe it identically each time.
Step 5: Batch generation — but test first
The original creator batch-processed prompts in Google Flow (16:9, one image per prompt, Nano Banana 2 model) using a browser extension that feeds prompts in one at a time — turning 103 rounds of copy-paste-click into one background job. 102 of 103 images generated successfully.
The critical detail: he tested one prompt first. Imagine discovering at frame 80 that your character changes weight every scene. Generate 2–3 test images, verify the style, then batch. And keep alternatives ready — any batch pipeline will have a few failures to redo manually.
Step 6: Assembly — timestamps get you 90%, judgment gets the rest
Drop the voiceover on the timeline, place images in order, trim each to the next timestamp. Then play it back and actually watch it.
Timestamps are a map, not a mandate. Some cuts feel too fast even when they’re “correct.” Information-dense frames deserve an extra second. A sudden turn in the narration sometimes lands better with an early cut. The last 10% is human judgment, and it’s the 10% that separates watchable from robotic.
Finally, back to Claude for title, description, tags, and 5 thumbnail concepts. Batch-generate thumbnails, pick the one with the clearest subject and sharpest tension, publish.
What the Workflow Doesn’t Tell You (The Parts That Actually Decide Success)
Here’s where I’ll add what the original analysis — and the hype video — leave out:
1. This is a survivorship story. For every faceless channel that hit 137K subs in two months, thousands more publish 14 videos into the void. The workflow explains how to produce the video. It says nothing about why this channel’s topics, titles, and thumbnails won — and that’s where the results actually came from. The production pipeline is table stakes; the topic selection is the edge.
2. YouTube is actively tightening policy on this exact content. YouTube now requires creators to disclose realistic AI-generated content, and its systems have gotten notably harsher on mass-produced, low-differentiation videos. When thousands of creators run the same tools, the same voices, and the same stick-figure style, the format commoditizes fast — and the algorithm’s tolerance for it shrinks. The original article hints at this; I’ll say it flatly: the window for generic faceless AI content is closing, not opening. Differentiation (a real angle, real data, a genuine voice) is what survives.
3. The monetization math needs revenue per view, not views. $77K/month on ~7M monthly views means roughly $10+ RPM across all sources — achievable in finance/business niches with sponsors or affiliate offers, but not from AdSense alone on most topics. If you’re modeling this, model the full revenue stack, and remember the video itself sold a workflow.
4. Start embarrassingly small. Don’t begin with 103 scenes. One specific topic, one 3–5 minute video, and measure three numbers: CTR (title/thumbnail), 30-second retention (hook), and average view duration (pacing). Each maps to a specific fix. A handful of these micro-tests will teach you more than any thread about someone else’s success.
The Honest Takeaway
AI genuinely solved something real here. Producing this video used to require a copywriter, voice artist, designer, and editor. Now one person can chain all four roles together in a weekend. That change is enormous, and it’s permanent.
But notice what AI didn’t solve: whether anyone watches, whether the format stands out, whether the channel compounds. Those still come from topic selection, hooks, and iteration — the unglamorous parts that no prompt can generate.
The tools are now a commodity. Everyone can download them.
The people who win are the ones who find subjects worth explaining, write scripts people actually finish, and run this loop enough times to get good at it. The workflow above is how you produce. What you produce — and whether anyone cares — is still on you.
Source & credit: This article is a rewritten and expanded adaptation of “14条视频涨粉13.7万:Claude Code+YouTube=月入7.7万美元” by 差不多小姐, published on the WeChat Official Account 小小创业经 (September 15, 2026): https://mp.weixin.qq.com/s/jfpJVYHP2vxcz1wponGh2Q — with additional analysis on monetization, platform policy, and survivorship bias. All views expressed in this adaptation are my own.