DEV Community

Cover image for I Made Spider-Man Swing Without Animating a Single Frame
Athreya aka Maneshwar
Athreya aka Maneshwar

Posted on AI-assisted

I Made Spider-Man Swing Without Animating a Single Frame

Hello, I'm Maneshwar, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.

Yes, that is Miles Morales.

He drops in on a web, lands a superhero crouch, fires a web off screen, and yanks a giant DEV board into frame.

I did not animate a single frame of that by hand.

I also had never opened Blender before this week.

This post is the story of how that happened, and the surprising amount of computer science hiding inside a "simple" 3D character.

Why Blender, of all things

I was trying to make a video.

Not a screen recording, an actual video with something moving in it.

My first stop was HyperFrames, which lets you write a video as HTML and render it out.

It is genuinely great at motion graphics.

Text that slides in, logos that pop, charts that draw themselves, all of that is lovely.

But the moment I wanted a character to do something, it hit a wall.

HTML can move rectangles around beautifully.

It cannot make a guy swing on a web.

So the next question was obvious, if a little scary: what do actual animators use?

The answer, for free and open source, is Blender.

And Blender has an MCP server, which meant I would not be learning it alone.

The MCP that drove the whole thing

The thing that made this possible for a total beginner is blender-mcp.

It connects Blender to an LLM client like Claude Code, so instead of hunting through 400 menus, you say "load this model and relink its textures" and it happens.

The architecture is worth understanding, because it is simpler than it sounds.

flowchart LR
    ME[Me typing a prompt] --> CC[Claude Code]
    CC -->|MCP over stdio| SRV[blender-mcp server]
    SRV -->|JSON over TCP socket| ADDON[Blender add-on]
    ADDON -->|runs Python| BPY[bpy API]
    BPY --> SCENE[Live Blender scene]
    SCENE -->|viewport screenshot| CC

    classDef human fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef agent fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
    classDef bridge fill:#6ea8ff,stroke:#2f62c4,color:#1a1a1a
    classDef blender fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a

    class ME human
    class CC agent
    class SRV,ADDON bridge
    class BPY,SCENE blender

There are two halves.

An add-on inside Blender opens a small socket server and waits for commands.

An MCP server outside Blender speaks the Model Context Protocol to the LLM on one side, and forwards JSON commands down that socket on the other.

The most powerful tool it exposes is execute_blender_code, which runs arbitrary Python inside Blender through bpy, Blender's Python API.

Every button in Blender is just a thin skin over bpy.

So an agent that can write bpy can, in principle, press every button.

The loop that makes it work is the screenshot tool coming back the other way.

The agent changes something, grabs the viewport, looks at it, and decides what to fix next.

That is the same edit, run, inspect loop you use on code, just pointed at a 3D scene.

One honest warning: that tool is arbitrary code execution by design. Treat it like you would treat any local automation with a shell, and do not point it at a machine full of secrets.

Step one: get a character

I am not a 3D modeler, and I was not about to become one by Friday.

So I grabbed a ready made Miles Morales model, the Spider-Verse Spider-Man, from Sketchfab.

Loading Miles into Blender for the first time

Even "just load it" had lessons in it.

The file came with four copies of Miles standing in a row: masked and unmasked, posed and T-posed.

The textures pointed at the original author's Windows folders, like C:\Users\someone\Desktop\..., so the agent had to relink them to the files in the zip.

And the zip was missing the normal and specular maps.

In Blender, a texture node pointing at a missing file renders hot pink.

It is the 3D equivalent of a 404, and it is very, very visible.

The fix was to disconnect those slots, at the cost of some surface detail on the suit.

Then I opened the unmasked Miles and his face looked like a horror movie, covered in white lines and dots.

Turns out the face was fine.

That model ships with a 143 bone face rig (eyelids, brows, lips, jaw, cheeks, even a tongue), and the armature was set to "draw in front", so every one of those bones was painted on top of his skin.

One checkbox later, he was handsome again.

The concept that unlocks everything: rigging

Here is where I learned the word that the rest of this post hangs on.

A 3D character is, at its core, a mesh: a big list of vertices (Miles has about 27k faces) glued into triangles.

A mesh is a statue.

You cannot make a statue wave.

To move it, you need three more things.

An armature, which is a skeleton.

It is a tree of bones, each one a parent of the next: hips to spine to chest to shoulder to upper arm to forearm to hand to fingers.

Rotate the shoulder and everything below it in the tree comes along for the ride, exactly like a transform hierarchy in any game engine or a nested <div> with a CSS transform.

Skin weights, which bind the mesh to the skeleton.

Every vertex stores a small table saying how much each bone influences it.

A vertex on the forearm might be 100% forearm.

A vertex right at the elbow might be 50% upper_arm and 50% forearm, which is what lets the elbow bend smoothly instead of folding like a paper straw.

The math is called linear blend skinning, and it is just a weighted average of matrices:

final_position(v) = Σ  weight(v, bone) × bone_matrix × rest_position(v)
                   bones
Enter fullscreen mode Exit fullscreen mode

A rest pose, which is the reference everything is measured from.

That is why every character you download stands in a T-pose, arms straight out.

It is not a dance move.

It is the zero point, the git init of the body, and every animation is stored as rotations relative to it.

Put those together and you have a rig.

And then the naive plan fell apart, because the model's rig had bones but no controls.

Gru's plan meme: download a Spider-Man model, make him dance, he has bones but no rig

Forward vs inverse kinematics

I had seen YouTube videos where an animator grabs a character's hips, drags them down, and the knees bend while the feet stay glued to the floor.

When I dragged Miles's hips, his whole body floated down through the floor like a ghost.

The difference is one of the nicest ideas in this whole field.

Forward kinematics (FK) is what a raw skeleton does.

You set the angle of every joint, and the hand ends up wherever the chain of rotations puts it.

Easy to compute, miserable to animate. Want the hand on the doorknob? Good luck guessing the shoulder, elbow and wrist angles that land it there.

Inverse kinematics (IK) flips the question.

You say where the hand (or foot) should be, and a solver works backwards to find joint angles that put it there.

It is literally an optimisation problem solved every frame, and Blender ships an IK constraint that does it for you.

Pin the feet with IK targets, drag the hips down, and the solver has to bend the knees because the feet are not allowed to move.

That is the magic in those YouTube videos.

It is not magic, it is a constraint solver.

Building a full IK control rig by hand is a real skill though.

And I wanted Miles doing backflips this week, not next quarter.

Mixamo: rigging as a service

Enter Mixamo, Adobe's free auto-rigger and animation library.

The flow is: export your character as an FBX in T-pose, upload it, and Mixamo builds a skeleton and skin weights for you.

The only thing it asks of you is to place a few markers.

Mixamo's auto-rigger asking for markers on the chin, wrists, elbows, knees and groin

Chin, wrists, elbows, knees, groin.

Five landmarks, with symmetry on, so really just three clicks per side.

From those it infers the whole skeleton, which is a nice reminder of how much structure a human body has. If you know where the joints are, the rest is mostly proportions.

Placement matters more than you would think.

My first pass had the wrists sitting on the forearms, which would have made his hands bend in the wrong place every time he shot a web.

I picked Standard Skeleton (65), which includes finger bones, because a two handed pull without fingers looks like a man wearing oven mitts.

And then it just... worked.

Miles rigged by Mixamo, playing an animation

The rigged character goes straight into Mixamo's library of motion captured animations.

These are real performers in mocap suits, recorded and cleaned up, so they have the weight shifts and tiny wobbles that hand keyed animation struggles to fake.

Jumps, climbs, hangs, throws, dances, falls, the lot.

Miles doing a jump to hang animation

Miles doing an attack animation

There was one trade off, and it hurt a little.

Mixamo builds its own body skeleton.

It does not know or care about the beautiful 143 bone face rig the model came with.

"With skin" and "without skin"

When you download an animation, Mixamo asks whether you want it with skin or without skin.

This is a good little lesson in separating data from presentation.

With skin gives you the mesh, the skeleton and the motion.

Without skin gives you only the skeleton and the motion, a few hundred kilobytes of bone rotations over time.

Because every clip targets the same 65 bone skeleton with the same bone names, you only need the skin once.

Everything else is just keyframe data keyed by bone name, and Blender happily applies it to any armature whose bones match.

So I downloaded the first clip with skin and the rest without.

And I downloaded a lot of them.

A Downloads folder full of Mixamo FBX files

Seventy eight files.

Clicked one at a time, because Mixamo has no multi select.

Some of them even share a name with a different animation inside, because the site just appends (1) to whatever you downloaded last.

They're the same picture meme: Hanging Idle.fbx vs Hanging Idle (1).fbx, Mixamo says they're the same picture

They are not the same picture.

The agent ended up scanning every FBX to see what was actually inside before trusting a single filename.

From a pile of clips to a performance

Here is the whole pipeline, end to end:

flowchart TD
    A[Sketchfab model] --> B[Clean up in Blender<br/>relink textures, drop pink slots]
    B --> C[Export T-pose FBX]
    C --> D[Mixamo auto-rig<br/>5 markers]
    D --> E[Download clips<br/>1 with skin, rest without]
    E --> F[Import as Blender actions]
    F --> G{Clip looks right<br/>from this angle?}
    G -->|no| E
    G -->|yes| H[Chain clips on NLA<br/>crossfades + root motion]
    H --> I[Webs, board, camera]
    I --> J[360p preview render]
    J --> K{Good enough?}
    K -->|no| H
    K -->|yes| L[Final render]

    classDef decision fill:#f4d35e,stroke:#b8991f,color:#1a1a1a
    classDef source fill:#e9ecef,stroke:#6c757d,color:#1a1a1a
    classDef rig fill:#5ee6c8,stroke:#1f9c86,color:#1a1a1a
    classDef scene fill:#9d8cff,stroke:#5b4bcc,color:#1a1a1a
    classDef out fill:#ff9a5c,stroke:#c4602a,color:#1a1a1a

    class G,K decision
    class A,B,C source
    class D,E,F rig
    class H,I scene
    class J,L out

Each Mixamo clip lands in Blender as an action, a named bundle of animation curves (an f-curve per bone per rotation channel).

On its own, an action is a single move.

A performance is several of them in a row, and that is what Blender's Nonlinear Animation editor is for.

The NLA is basically a video editor timeline for animation.

Each action becomes a strip, strips sit on tracks, and overlapping strips can blend into each other.

That blend is what stops a cut from "hanging" to "landing" from snapping like a broken GIF.

The agent wrapped all of this into a tiny sequencer, so the shot is described as a list of moves instead of a pile of frame numbers:

perf = Performer(rig)
perf.then("Hanging Idle", length=40)
perf.then("Hard Landing", to=50, speed=1.3, blend=6)
perf.then("Standing 1H Magic Attack 02", frm=16, to=40, speed=1.25, blend=6)
perf.then("Pull Heavy Object Stop", to=30, speed=1.1, blend=5, face=65)
perf.then("Victory Idle", blend=8, face=10, length=51)
perf.build()   # NLA strips, crossfades, stitched root motion
Enter fullscreen mode Exit fullscreen mode

blend=6 is a six frame crossfade into the next clip.

face=65 turns him 65 degrees so the yank points at the board.

And the "web shot" is, I promise, a clip literally named Magic Attack. Mixamo does not have a "thwip".

The interesting problem hiding in build() is root motion.

Some clips move the character through space (a landing drops him, a pull leans him back), and each clip starts from its own origin.

Play two of them back to back naively and Miles teleports back to the start between moves.

So the sequencer reads where the hips ended in one clip and offsets the next one to begin there.

It is the same idea as stitching GPS tracks together: each segment is relative, so you keep a running sum.

Because all timing derives from that list, swapping one clip re-times the webs, the board and the camera automatically.

The web strands, the board and the camera are not Mixamo at all, they are plain Blender objects keyframed by script to fire on the frame where his hand reaches out.

Iterate at 360p, like a sane person

The first full render was 16 seconds at 720p and took about eight minutes on my GTX 1650.

That is fine once.

It is not fine when the feedback is "make it snappier", "make the board bigger" and "his suit is too dark" in the same breath.

So previews dropped to 360p, roughly 4x faster, and 720p became a finals only thing.

The same rule you know from code applies: make the feedback loop fast first, then make the output good.

Two more things that cost real time:

The suit rendered as a flat black blob.

The model used a toon shader that collapses lighting into a couple of hard bands, which looks great in a cartoon and terrible when you want to see muscles.

Swapping in physically based shading, plus a light just for him and red and blue rim lights, brought his shape back.

Light grey or dark studio?

A near black suit on a near black background is a silhouette of nothing.

Lighter backgrounds read better. Dark backgrounds look cooler. The GIF at the top is the dark one, because I have priorities.

Escaping the download button

After clicking "Download" 78 times, I asked the obvious engineer question: can I just do this locally?

Mixamo's rigger and library only live on Adobe's servers.

But motion capture data is not an Adobe invention.

The CMU Graphics Lab Motion Capture Database has about 2,500 free clips, and there is a BVH mirror on GitHub with a text index of what each one is.

Cartwheels, salsa, climbing, martial arts, "pulling a rope", all of it.

The catch is that CMU's skeleton is not Mixamo's skeleton.

Moving motion from one skeleton to another is called retargeting, and it is a lovely little linear algebra problem.

You cannot just copy rotations across, because the two skeletons have different rest poses. A 30 degree rotation from "arms slightly down" is not the same pose as 30 degrees from "arms straight out".

So the retargeter converts each bone's rotation out of the source rest pose, into a shared world frame, and back into the target's rest pose.

It also scales the hip travel by leg length, so a tall performer's stride does not make Miles moonwalk.

The first try rendered... nothing.

Miles was not in frame. He was not anywhere.

The BVH rest pose put the hips exactly at the origin, so "rest hip height" was zero, and the scale factor was a division by it.

Miles had been launched into orbit by a divide by zero.

Fixed by measuring the floor from the lowest foot in the clip instead, and suddenly he was walking, dancing and cartwheeling from a dataset recorded at Carnegie Mellon years ago.

What I actually learned

I set out to make a guy swing on a web.

I came back with a list of concepts I did not expect from an art tool:

  • A character is a mesh deformed by a tree of bones, via per vertex weights.
  • Animation is rotations relative to a rest pose, which is why everything T-poses.
  • IK is a constraint solver, and it is the difference between a puppet and a corpse.
  • Motion clips are data keyed by bone name, so one skin can wear a thousand animations.
  • NLA strips are a timeline of reusable actions, and chaining them means stitching root motion.
  • Retargeting is a change of basis, and it will divide by zero if you let it.
  • Blender is a Python program with a UI on top, which is why an MCP bridge can drive it at all.

There is a ton more in there for a CS person: shader node graphs are dataflow programs, Cycles is a path tracer, physics sims are numerical integrators, geometry nodes are basically a functional language for meshes.

But one step at a time.

Next on the list is the face rig I so cruelly abandoned in that pool, and maybe making Miles actually talk.

If you have played with Blender, rigging or the MCP, tell me what I should pick up next.



Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production secure and reliable without slowing you down.

I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.

Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.

Spend code review effort where business risk is highest — not spread evenly across every diff.

⭐ Star it on GitHub:

GitHub logo HexmosTech / LiveReview

Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview

gitleaks.yml osv-scanner.yml govulncheck.yml semgrep.yml dependabot-enabled mcp-testcases.yml

LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.

blast-radius-demo.mp4

LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
















The exact math, not a black box Visualize blast radius at a glance Every factor that feeds the score

How does Blast Radius scoring work? (a more technical explanation)

Here's the goal:

  • A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
  • A 300-line UI change in one file, fully covered by…




Click below to try LiveReview with your codebase:

LiveReview Banner

Top comments (3)

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥 •

Wow 🥰

Collapse
 
aidiveyt profile image
AI Dive •

Since 2.1.228 Claude Code starts Pro/Max/Team sessions in auto permission mode by default, so an execute_blender_code call can land with no prompt at all. Worth checking which of the six modes you are in before wiring an MCP that runs arbitrary code.

Collapse
 
t3ratech profile image
Teratech Solutions •

The MCP → Blender bridge is a great pattern. One gotcha we hit: MCP tools that mutate scene state (add keyframes, modify constraints) need idempotency keys — our agent retried a "parent bone to mesh" call after a network stall and ended up with the rig parented twice, which corrupted the export. Idempotent mutations + a dry-run tool mode made the whole lane reliable. Curious if you hit the same class of failure when the agent replays a turn.