Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.
Two posts ago, I had never opened Blender.
Since then, Miles Morales has swung on a web, yanked a DEV board into frame, and moonwalked 6.5 metres off the edge of a billboard without anyone noticing.
Every one of those shots was me and Claude Code, pair programming a 3D scene one prompt at a time.
It works.
It is also slow, because I am still the one holding the whole movie in my head.
So this post is about the next question, the one I keep circling back to at 1 a.m.:
What would it actually take for me to type one paragraph and get a finished YouTube video out the other end?
Not "Claude helps me in Blender".
Claude directs the thing.
I went digging through every MCP server, Claude Code skill, mocap source and audio model I could find, checked the GitHub stars so we know what is battle tested and what is a weekend project, and wrote down the stack I would build.
Spoiler: Blender is maybe a third of it.
The job description: director, not animator
The single most useful thing I learned in part one was not about rigging.
It was about where the LLM is good and where it is hopeless.
Claude was great at writing bpy code, relinking broken textures, sequencing clips on a timeline, and stitching root motion between them.
Claude was never once asked to invent how a human body lands from a jump.
That came from Mixamo, which came from a real performer in a mocap suit.
And that split is the whole architecture.
An LLM is a brilliant writer of structure and code, a decent critic of a still image, and a terrible source of biomechanics.
So every part of a video should go to whoever is actually good at it:
If it moves like a human, fetch it.
If it is code, write it.
If it is a sound, order it from a model that does nothing but sound.
Here is the whole stack with that rule applied:
flowchart TD
ME[One paragraph from me] --> DIR[Claude Code<br/>director skill]
DIR --> SL[Shot list<br/>YAML]
SL --> BR[Blender bridge<br/>one MCP server]
SL --> MO[Motion library<br/>Mixamo, CMU, Rokoko]
SL --> AS[Assets<br/>Poly Haven, Sketchfab, image to 3D]
SL --> AU[Audio<br/>TTS, SFX, music]
AU --> LS[Rhubarb<br/>mouth shapes]
MO --> BR
AS --> BR
LS --> BR
BR --> BL[Blender, headless<br/>EEVEE drafts, frame sequences]
BL --> QA{Checks pass?<br/>eyes + asserts}
QA -->|no| DIR
QA -->|yes| FF[FFmpeg<br/>frames + stems]
AU --> FF
FF --> YT[YouTube-ready MP4]
classDef human fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
classDef brain fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
classDef fetch fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
classDef order fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a
classDef engine fill:#6ea8ff,stroke:#2f62c4,color:#1a1a1a
classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
class ME,YT human
class DIR,SL brain
class MO,AS,LS fetch
class AU order
class BR,BL,FF engine
class QA decision
Let's walk it bottom up, starting with the piece everyone argues about.
Layer 1: the Blender bridge (pick exactly one)
This is the MCP server that lets Claude reach into a running Blender and press buttons by writing Python.
There are now a lot of them.
Here are the four I looked at seriously, with stars as of this week:
| Bridge | What it is good at |
|---|---|
mcp-for-blender ⭐ (~30.2k) (formerly blender-mcp) |
The one I used in parts 1 and 2. Code execution, viewport screenshots, plus Poly Haven, Sketchfab, Poly Pizza, Hyper3D Rodin, Hunyuan3D and Tripo built in |
| Official Blender MCP ⭐ by Blender Lab | Understanding a scene: explaining node setups, finding what uses a material, cleanup. Blender's page lists 5.1 or newer |
| carlosh7/blender-mcp ⭐ (8 stars) | 248 tools across 14 categories, 371 commits, 479 passing tests, MIT |
| claude-blender ⭐ (5 stars) | Undo checkpoints, camera auto-framing, local TripoSR and Shap-E generation |
The official one is official in the sense that Blender developers built it, and Anthropic announced it as a Claude connector this year.
Read its own page though, and it is pitched as a way to understand complex setups, not to churn out a film.
It also carries the most honest warning I have seen on any MCP page: the server "will execute LLM generated code in Blender without any guards in place", and Blender suggests a VM or a machine without sensitive data on it.
That applies to every server in the table, by the way.
They all ship some flavour of "run this Python", and that is the tool doing the real work.
Which brings me to the 248 tool one.
It is clearly a lot of careful work, with real tests and CI.
But it made me realise that tool count is not a feature.
Every tool an MCP server exposes comes with a schema: a name, a description, a JSON shape for its arguments.
Many clients put all of those in the context window on every single turn, whether you use them or not.
Some, Claude Code included, can defer them and search on demand, which fixes the space problem.
It does not fix the picking problem.
If there are five slightly different ways to move a cube, the model has to choose one, and every choice is a chance to choose wrong.
One execute_blender_code and a model that writes good bpy already covers all 248, because every button in Blender is a thin skin over bpy anyway.
So my pick is boring: mcp-for-blender as the one bridge, installed with claude mcp add blender uvx mcp-for-blender, and the official one kept around for "why is this scene slow" archaeology.
And I am stealing the best idea from the 5 star repo without installing it.
Checkpoints before risky edits are a great idea.
In part two the .blend file became a build artifact generated by scripts, so my checkpoints are just git commit.
Layer 2: skills, not more tools
If tools are expensive, how do you teach Claude cinematography?
With Claude Code skills.
A skill is a folder in .claude/skills/ with a SKILL.md inside: a one line description that is always visible, and a body of instructions and scripts that only loads when the task needs it.
It is the lazy import of prompting.
blender-skills (~143 stars) is a good example of the idea.
It has camera moves as skills, turntable, slow-zoom, dolly-rotate, crane-shot, perfect-loop, plus image to 3D through Meshy and a toolkit that does Mixamo retargeting.
It is aimed at product shots, so I would not install it wholesale for a superhero short.
I would copy its shape.
The thing I actually want is a small library of film vocabulary, each word backed by tested code:
-
establishing_wide,push_in,pull_out,orbit,crane_up,tracking,over_the_shoulder,handheld_shake -
rainy_night,golden_hour,villain_red,neon_street,soft_closeup_key
Then "start wide, track him to the ledge, push in slowly as he says the line" stops being a creative writing exercise for the model.
It becomes three skill calls with numbers.
Same idea applies to lighting.
A preset that is right once is right forever, and nobody has to rediscover three point lighting at 2 a.m. because Miles's suit rendered as a black blob again.
Layer 3: motion, fetched not generated
This is where most "AI made my animation" demos fall apart, and it is the layer I am most opinionated about.
Do not let the LLM keyframe a human.
It will produce something that technically moves and spiritually does not.
Real motion has weight shifts, overshoot, a little settle at the end of a landing.
Mocap has all of that for free, because a person did it.
The sources I would wire in, in order of effort:
Mixamo. Free, auto-rigs a T-posed character from five markers, thousands of clips. I downloaded 78 of them one click at a time in part one, and I would do it again.
CMU Motion Capture Database. About 2,500 free clips, with a BVH mirror on GitHub (~203 stars). Needs a retargeter, which part one built and part two debugged.
Rokoko. Their Blender add-on is free and does retargeting. Their free plan includes Rokoko Vision, single camera AI mocap. That last one is the exciting bit: the motion Mixamo does not have, I can act out in my living room with a webcam.
And then the controls.
In part one I dragged Miles's hips down and his whole body sank through the floor like a ghost, because the rig had bones but no IK.
Two fixes, depending on budget:
Mixamo Rig is a free, GPL extension that builds a proper IK control rig on top of a Mixamo skeleton and bakes animation in and out of it. Check the version range on the listing before you build on it.
Auto-Rig Pro by Artell is the paid, serious option, about $50, used in shipped games like Manor Lords, with rigging, retargeting between skeletons and facial rigs in one place.
If I spend money anywhere in this stack, it is here.
Rigging is the one job where a mistake quietly poisons every shot that comes after it.
Layer 4: the face, or the close-up problem
Here is a thing I did not appreciate until I tried to frame a close-up.
In a wide shot, nobody looks at the face.
Then the camera pushes in for the line, and the face is the only thing on screen.
A mouth flapping open and closed on a timer reads instantly as fake.
So does a character who never blinks, or whose eyes are locked dead ahead like a mannequin's.
The good news is that most of facial performance can be fetched too, and the most important piece is wonderfully small.
Rhubarb Lip Sync (~2.6k stars) is a command line tool that listens to a voice recording and writes out a timeline of mouth shapes.
It uses nine shapes, A to H plus X for rest, the same small set classic cartoon studios used, and you can hand it the script with --dialogFile so it knows the words instead of guessing.
It says it is built for 2D, but the output is just JSON with timestamps, and a 3D face can use it just as well.
There is a Blender add-on for it (~228 stars), but it has not been touched since 2023.
The glue is short enough that I would rather have Claude own it:
# rhubarb -f json -d line.txt -o line.json line.wav
import bpy, json
cues = json.load(open("line.json"))["mouthCues"]
keys = bpy.data.objects["Miles_Head"].data.shape_keys.key_blocks
fps, start = bpy.context.scene.render.fps, 120 # line begins on frame 120
# Rhubarb's letters -> this model's shape key names
VISEME = {"A": "MBP", "B": "EE", "C": "EH", "D": "AA", "E": "OH",
"F": "OO", "G": "FV", "H": "L", "X": "rest"}
for cue in cues:
frame = start + round(cue["start"] * fps)
for shape in VISEME.values():
keys[shape].value = 1.0 if shape == VISEME[cue["value"]] else 0.0
keys[shape].keyframe_insert("value", frame=frame)
That is the whole trick: audio in, a keyframe per mouth change out.
Blender's default Bezier interpolation eases between the keys, so the mouth glides instead of snapping.
If your model drives its mouth with face bones instead of shape keys, like the 143 bone face rig on my Miles, the same loop poses bones instead.
Blinks and eye darts are even cheaper.
A blink every two to six seconds, randomised, and a tiny eye flick toward whatever the character is reacting to, is a dozen lines of procedural keyframing, and it is the difference between a person and a wax figure.
Notice the dependency hiding here, too.
Lip sync needs the audio to exist first.
Which means the voice is not post production.
It is pre production.
Layer 5: assets, mostly fetched
Same rule again: generate only what nobody has made yet.
Poly Haven. Thousands of CC0 HDRIs, textures and models. CC0 means no attribution and no strings. mcp-for-blender searches and imports from it directly, and an HDRI alone fixes half of "this scene looks fake".
Sketchfab. Huge, and also where my Miles came from. Read every license, because they vary model to model.
Image or text to 3D, for the props nobody has made:
- Hunyuan3D-2 (~15k stars) and Hunyuan3D-2.1 (~4.1k stars, adds PBR materials), from Tencent, runnable locally if your GPU is braver than my 4 GB GTX 1650
- TripoSR (~7k stars), fast single image reconstruction
- Hosted options like Hyper3D Rodin and Tripo, already wired into mcp-for-blender
One honest caveat from someone who just learned what skin weights are.
Generated meshes are usually a dense soup of triangles.
That is fine for a mailbox in the background.
It is not fine for anything that has to bend, because rigging wants clean edge loops around elbows and knees, and a triangle soup deforms like wet cardboard.
So: generated props, sourced or hand built heroes.
Layer 6: audio is half the movie
Every engineer I know who tries video makes the same mistake I did.
They spend three days on render settings and zero minutes on sound.
Put a whoosh on a jump, a thwip on a web shot and a low thud on the landing, and the exact same frames suddenly look like they cost ten times more.
The trick is to treat audio as separate stems, never one mixed track:
Dialogue. ElevenLabs for character voices. Generate it first, because the face needs it.
SFX and ambience. ElevenLabs sound effects turns a text description into a sound, up to 30 seconds per generation, with a seamless loop mode for rain and city hum, and 48 kHz WAV for one shots.
Music. Its own model, its own stem. stable-audio-tools (~3.9k stars) is Stability's open code for their audio models if you want local control. Check the license of whatever you pick before it goes on a monetised channel.
Separate stems are what make the edit programmable.
Music too loud under the line? Change one number.
Want the music to dip automatically whenever someone talks? That is FFmpeg's sidechaincompress filter, keyed off the dialogue stem.
None of that is possible once everything is baked into one WAV.
Layer 7: render like a build system, assemble with FFmpeg
Part two already turned the scene into a build artifact, so this layer is mostly "keep doing that, at movie scale".
Three rules:
Render headless. blender -b with no UI, driven by scripts, so a render is a command, not a person clicking.
Render frame sequences, not MP4s. If Blender crashes on frame 1,400 of a movie file, you lose the movie. If it crashes on frame 1,400 of a PNG sequence, you lose one frame and resume. Blender's own docs recommend exactly this.
Draft at 360p, final once. Part two's --fast mode, with ray tracing, motion blur and extra samples switched off, roughly halved frame time. Everything is reviewed in drafts.
Then FFmpeg (~64.9k stars, and it is only a mirror) glues frames and stems into a video:
# render shot 30 as a PNG sequence (flags before -a, Blender reads them in order)
blender -b build/sh030.blend -o //frames/sh030_#### -F PNG -s 1 -e 96 -a
# frames + three stems -> one mp4, music ducked under the dialogue
ffmpeg -framerate 30 -i frames/sh030_%04d.png \
-i audio/dialogue.wav -i audio/sfx.wav -i audio/music.wav \
-filter_complex "[1:a]asplit[d][key];[3:a][key]sidechaincompress=threshold=0.05:ratio=8[m];\
[d][2:a][m]amix=inputs=3:normalize=0[a]" \
-map 0:v -map "[a]" -c:v libx264 -pix_fmt yuv420p -crf 18 -c:a aac sh030.mp4
asplit makes two copies of the dialogue, one to hear and one to act as the key that pushes the music down.
normalize=0 stops amix from quietly turning everything down to make room, which is a fun one to debug by ear.
-pix_fmt yuv420p is there because some players will happily show a black screen without it.
I would still do the final colour pass and titles in DaVinci Resolve for anything going on my channel, but a machine should be able to produce the rough cut alone.
Layer 8: two reviewers, because one of them is blind
The loop that made parts one and two work was Claude grabbing a screenshot, looking at it and fixing what it saw.
That loop is great.
It is also how a moonwalk got shipped.
Vision models review a still picture.
Motion bugs live between the frames.
The 22 cm pop in the middle of a crossfade, the single frame where Miles hovered at a cut, the stance wider than the ledge he stood on, none of those survive a look at frame 1,201.
They all fail an assertion that measures where the feet are on every frame.
So every shot gets both reviewers:
- Eyes: Claude looks at a contact sheet of draft frames and judges framing, lighting, mood and readability.
- Asserts: scripts measure contact, bounds, continuity at cuts and missing textures, and fail loudly.
If the asserts fail, nobody wastes time looking.
The missing piece: a director skill
Everything above is a department.
What turns departments into a studio is one skill that takes my paragraph and writes the plan everyone else works from.
That plan should be data, not prose:
title: rooftop_siren
fps: 30
character: original_hero_v2 # not Miles, see above
shots:
- id: SH010
seconds: 4
set: rooftop_night_rain
camera: [establishing_wide, push_in]
motion: { clip: "Standing Idle" }
audio: { ambience: "rain on a rooftop, distant traffic, wind" }
- id: SH020
seconds: 2.5
camera: [extreme_closeup]
motion: { clip: "Look Around", turn: 30 }
dialogue: { voice: hero, line: "You hear that?" }
face: { lipsync: rhubarb, blinks: auto, look_at: siren }
- id: SH030
seconds: 3.2
camera: [tracking, low_angle]
motion: { clip: "Jump Down", speed: 1.2 }
audio: { sfx: ["cloth whoosh", "web thwip", "heavy landing on concrete"] }
Look at what is in there.
Every value is a skill name, a library lookup or a prompt for an audio model.
Nothing asks the LLM to invent a joint angle, and nothing asks it to remember the whole movie at once.
The director writes this once.
Then each shot runs through the same loop, and a bad shot is re-run on its own instead of the whole film:
flowchart LR
S[Shot from the YAML] --> A[Generate dialogue<br/>and SFX]
A --> L[Rhubarb<br/>mouth cues]
S --> M[Fetch motion<br/>and set]
L --> B[Build shot<br/>bpy scripts]
M --> B
B --> D[360p draft]
D --> T{Asserts pass?}
T -->|no| B
T -->|yes| V{Claude likes<br/>the contact sheet?}
V -->|no, fix camera or light| B
V -->|yes| F[Final frames]
F --> X[FFmpeg mux]
classDef src fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
classDef order fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a
classDef fetch fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
classDef build fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
classDef out fill:#6ea8ff,stroke:#2f62c4,color:#1a1a1a
class S src
class A order
class L,M fetch
class B,D build
class T,V decision
class F,X out
It is CI for a movie.
Each shot is a job, the YAML is the config, the asserts are the tests, and Claude is the reviewer who also writes the fixes.
The order I am building it in
Not the order of the diagram, the order of the dependencies:
- Bridge + skills. One MCP server, a camera and lighting skill library. Without reliable shots, nothing else matters.
- Motion library. Mixamo and CMU behind one retargeter, Mixamo Rig for IK, Rokoko Vision for the clips nobody has.
- Audio stems. Voice, SFX, music, all separate. This comes before the face, because the face reads the audio.
- Face. Rhubarb glue, procedural blinks and eye targets, then close-ups stop being scary.
- Asserts. Part two's checks, generalised to every shot.
- Director skill. Only once every department works on its own. A director with no crew just writes very confident YAML.
And the question I will be honest about: will this make a Pixar short?
No.
What it can make is something I could not make at all a month ago: a 30 second scene with a character who moves like a person, talks with a mouth that matches the words, sounds like it happened somewhere, and was assembled by a machine while I wrote the next paragraph.
That is a wrap for part three.
If you have built any piece of this, especially the face, tell me what broke. I clearly have a lot of breaking ahead of me.
Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.
I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.
Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.
Spend code review effort where business risk is highest — not spread evenly across every diff.
⭐ Star it on GitHub:
HexmosTech
/
LiveReview
Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems
LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.
blast-radius-demo.mp4
LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
Here's the goal:
- A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
- A 300-line UI change in one file, fully covered by…
Click below to try LiveReview with your codebase:












Top comments (0)