After the kind response to my first post, where I wrote about building three websites in August, here's what came next. In September I started making music videos for the songs I create with AI and publish on my YouTube channel.
Let me be honest right away: all of this is play and experimenting. I'm still a beginner. I'm not a graphics programmer or a video editor. I just want to see how far I can go, because AI tools today have almost no limits.
In one month I made 45 music videos. 24 are the classic kind: scenes generated with AI video tools, then put together in a video editor. 21 are made entirely in code, with no video footage and no AI video.
I don't write the code myself. AI does. I come up with the story and the look, watch every version and say what's wrong. The AI writes and fixes the code.
Why code, when AI video exists?
- It's more fun for me. I get to watch something being built from zero, piece by piece.
- It looks different. It's not yet another "wow AI animation" like the thousands already online.
- I have full control. Code is something tangible: I can change any detail exactly the way I want. With pure AI video, the result always drifts off in its own direction, no matter if the prompt is short, long, detailed or vague.
How the work goes
The process was the same for all three videos:
- I write a script for the song. Line by line, with timestamps: at 0:41 this happens, at 1:24 that. This turned out to be the most important step.
- The AI writes the code and tests it on its own server. It doesn't render the whole video, just 20–30 still frames from different parts of the song, like a contact sheet.
- I look at those stills and say what's wrong. Most mistakes are visible in a picture, not in the code.
- When the stills look right, I render the real video on my own device, usually my phone, and report what still looks wrong in motion.
- Repeat until it's right.
I picked three videos. The first two use the same tool, Three.js in the browser, but each one solved a different problem: a city that follows the song and its lyrics, and a character that walks and acts. The third one is made in a completely different way, with Python.
1. Veins of the City (Three.js, in the browser)
A walk through a rainy neon city at night, where the lyrics light up as neon signs. Both the song and the video were inspired by Wong Kar-wai's film Fallen Angels.
How it was built, step by step:
- First, the song gets measured. When you pick the mp3, the browser (Web Audio) isolates the bass and finds the beats, about 250 in 5 minutes. That's a list of timestamps, and from then on the picture "knows" where the rhythm is.
- The lyrics become a table. Each line has three values: the text, when it lights up and when it goes dark. The auto-generated subtitles were a mess, so the timing was put together from my lyrics and the parts of the subtitles that were actually right.
- The song decides the shape of the city. The camera walks faster in the chorus and slower in the quiet part. From that speed, the length of each street is calculated, so the camera turns a corner exactly when the song changes sections. Song first, city second.
- Everything is drawn by one function: "show me moment t". You give it a time in the song, and it places the camera, lights the right signs and windows, and draws the frame. There's no real randomness: even the "random" flickering is calculated from the time, so it's identical every time.
- The frame goes through filters: glow, motion trails, blur toward the edges, then color grading and film grain.
- Export: the same function is called for moment 0, then 1/30 of a second, then 2/30… Each frame goes into an encoder (WebCodecs), which packs them together with the audio into an MP4.
Why step 4 matters most: when I say "something's wrong at 2:36", the AI can render exactly that frame and look at it. The same function also makes the vertical version for Shorts, just with a different frame shape.
Here's the idea in a few simplified lines:
// Everything on screen comes from one number: the time in the song.
function renderAt(t) {
moveCamera(t); // where we are on the street at this second
lightSigns(t); // which lyrics are glowing right now
flickerWindows(t); // windows react to the bass
drawFrame(); // draw the picture
}
// Export: ask for every moment, 30 times per second, and save each picture.
for (let frame = 0; frame < totalFrames; frame++) {
renderAt(frame / 30);
saveFrame();
}
What went wrong:
The server had no graphics card, so one frame took 3 seconds and the whole video about 4 hours. So the server only makes test stills, and I render the real video myself. On a desktop it took a few minutes.
The camera jerked after the second minute. The AI wrote a small program that "walks" the camera through the whole video and flags every spot where it jumps too much. It found the causes in seconds.
The camera walked into a building, and lyrics ended up in the wrong street. Both only showed up on the contact sheet.
2. Memories I Was Given (Three.js, in the browser)
A lonely robot walks through a burned-out world at sunset, with someone else's memories on its chest screen.
How do you get a 3D world without 3D software? There's no modeling program where you shape objects with a mouse. The robot is written as a list of parts: the head is a sphere, the visor is a ring, the chest is a rounded box. The parts are connected like joints: the hand hangs from the elbow, the elbow from the shoulder, the shoulder from the chest. When the shoulder turns, the whole arm follows. The ruins are made the same way: the ground is a grid of points whose heights are calculated, and a wall is a box with its top edge "bitten off". Three.js turns all of that into 3D shapes, and the graphics card draws them. The world is still 3D, it's just written instead of built by hand.
How it was built, step by step:
- A script, line by line. Before any code, the whole song was planned out: which shot, what the robot does, on which lyric. In the end this was the most useful part of the whole project.
- The song gets measured once, in advance. The finished video is a single HTML file, but during preparation Python was used once: it split the song into 10 frequency bands, 30 times per second, and the resulting numbers (about 50 KB) were written straight into the HTML. Python isn't needed to make the video, and the equalizer and the heart on the robot's chest always move the same way.
- The robot is built from parts, as described above, and the metal gets rust painted on a canvas.
- Walking is calculated backwards. I don't move the knee and hip by hand. I say where the foot should be, and the code calculates the angles (this is called IK). Each step lasts exactly one beat of the song.
- The whole path is calculated once, at load time: where the robot is and which foot is on the ground, for every moment of the song. That's why you can jump to any second and the feet still stay planted.
- Same idea as the city: "show me moment t". That function moves the sun, poses the robot, draws the chest screen and places the camera.
- A camera with no hard cuts. Within one location the camera glides from shot to shot. When the location changes, both shots are drawn and dissolve into each other.
- Export works like the city: frame by frame into an MP4.
What went wrong:
It worked on the AI's server, but on my computer (Windows) the screen was completely black. One broken pixel, which the glow effect smeared across the entire frame. The first fix turned the screen white, because Windows had quietly removed the part of the code meant to catch the problem. Lesson: test on another computer as soon as you have your first frame.
Screen recording didn't work: the video stuttered, and my video editor wouldn't even show the file. That's where the rule came from: always render frame by frame.
Vertical video for Shorts: a simple crop loses half the shot, so the vertical mode widens the field of view on its own.
Surprise: my mid-range phone exported the whole video in 1080p in about 7 minutes, right in the browser.
3. The Galaxy Devourer (Python, no browser)
A black hole devours a galaxy, styled as a Japanese arcade game from the 80s.
How it was built, step by step:
- Rhythm first. librosa finds the beats in the song. Since they "wobble" a little, the code lines them up into a perfectly even grid. From that come two small functions: "which beat are we on" and "how far into it are we". Everything else asks them.
- Then the rules of the look: a small image (384×216), only 16 colors, letters without smoothing. Those rules apply to every single frame.
- Every scene is a function of time. The black hole, the galaxy, the bonus stage, the captions: each one gets the time, asks "which beat is it", and draws its part (Pillow). Movement happens in steps on the beat, like in old games.
- One function for swallowing. Planets, dust, the galaxy and the player's ship all use the same function: "pull this point into the hole along a spiral". Written once, used everywhere.
- Captions and lyrics come last, on top of everything: the dialog box, the score counter, the stage name.
- Upscaling: the small image is enlarged five times (NumPy), then old-TV scanlines are added.
- No images on disk: each finished frame goes straight into ffmpeg, which joins it with the song. 6,100 frames in about 5 minutes on the server.
What I learned:
When everything hangs on the rhythm, the video feels in sync even when the animation is simple.
Limits help: low resolution and 16 colors created the style by themselves.
First a 30-second test, agree on the look, and only then the full video.
What applies to all three
- Almost everything on screen is calculated from the time in the song. If something's wrong at 2:36, you can render exactly that frame and look at it.
- Render frame by frame, don't record the screen. Recording stutters. This way it's identical every time, with no glitches.
- Look at frames, not code, but don't be afraid of the code. As a beginner I can't read 1,300 lines of code, but I can see that the camera is inside a wall. Sometimes I still ask to see the code and have the AI explain what each part does and how I can change it myself. That way I can tweak small things like a color or a speed on my own, and I learn something new every time.
- Test on a real device, and don't underestimate your phone. What works on one computer doesn't have to work on another. I only used my computer when I was impatient and wanted to render faster and see the result. Along the way I figured out that everything works on a phone too, so I did about 98% of the work on my phone.
Every video taught me something new, and every time I'm surprised how far you can get with AI, without any background in graphics or animation.
All the videos are on my YouTube channel: https://www.youtube.com/@kuga-i7o
Thanks for reading. U zdravlje! 😄
This post was written with help from AI, based on my own work.
Top comments (0)