Moving heads
A row above the frame and a row on the floor. Where each one points and how bright it is, is written for this song, bar by bar.
create-music-video · open-source Agent Skill
I Will Not Promote
1:29 · Sung in English · the hook is a line 344 different people had typed
It makes one kind of music video: a dark stage, lights that follow the band, every lyric in type as it is sung, and a drawn singer if you want one. Not generated footage, and not a story told in scenes.
Name an audience. The agent reads where they talk, builds the lyrics on phrases they really say, and has a music model sing them. Then it measures the song — its bars, who is playing, when each word is sung — and writes the stage against those measurements in Remotion. It renders on your machine.
Real output
This is the video the Skill was written from. We named an audience — solo founders who build and never promote. The first lyric the agent wrote, without reading anything, we turned down. So it read 1,849 posts from nine founder communities and counted: the hook is a line 344 different people had typed. We agreed the lyrics, chose the second of two performances, turned down four singers for being too good-looking, and picked one of the next four. Then the steps were written down as the Skill.
Every video the Skill makes is this kind. Another song gets other colours, other cuts and another singer, on the same stage.
Numbers from the video above: I Will Not Promote
“A song about my app” is a topic. “A song for solo founders who never promote what they build” has an audience, and the lyrics are going to come out of their mouths. Bring a song file of your own instead, and it starts at the measuring.
How many different people say a phrase in the same words, how recently, and what it means where it is said. One clever post is one person’s writing. A line three hundred people typed is a chorus.
Two or three hooks first, each with its count. Then the lyrics with the phrase under every line — before any song is made, because a lyric you do not like is not rescued by a good tune.
A music model sings the words as written, twice. The agent cannot hear either one. You listen to both, pick, and say whether anything is shouted that is not in the lyrics.
Three or four candidates that really differ, drawn and shown together, because nobody can choose a face from a description. Then every shot is drawn from the one you picked.
Bars, who is playing in each, when each word is sung. The lights and the cuts are written against that. It renders the frames that have to be right, looks at them, and fixes what is wrong before you see it.
// src/promote.phrases.json
{
"phrase": "I will not promote",
"people": 344,
"newest": "2026-09-16",
"where": "r/startups",
"means": "the line every post there has to carry in its title; people resent typing it"
} bar from kick voice and chords
13 18.26s ## ## ## ## ##### ##### ##### ####
14 19.72s # ###### ###### ###### #####
15 21.20s #### ### ### ## #### #### ##### #####// src/promote.lines.json
{
"section": "chorus 1",
"text": "I will not promote",
"how": "heard",
"by": "its own sound",
"seconds": [19.67, 20.11, 20.47, 20.83]
}A row above the frame and a row on the floor. Where each one points and how bright it is, is written for this song, bar by bar.
Colour thrown across the walls and floor, and haze drifting through it — which is what makes a beam visible at all.
One drawn person in eight to twelve shots. What makes them move is the cutting: each picture framed several ways, each cut landing on a bar or a word.
Told lines fill in along the bottom, word by word. Shouted lines are set as large as the frame will hold, and land a word at a time.
Sparks, paper, lasers, and one full-frame white flash. The second chorus has to have somewhere to go that the first did not.
A strip of meters moving to the song’s own spectrum, and a frame of readings that says what the lyric says in another voice.
$skill-installer install https://github.com/anglesvideo/angles-video-skills/tree/main/skills/create-music-videogit clone --depth 1 https://github.com/anglesvideo/angles-video-skills.git ~/angles-video-skills
mkdir -p ~/.claude/skills
ln -s ~/angles-video-skills/skills/create-music-video ~/.claude/skills/create-music-videoMake a ninety-second song for people who run a one-person newsletter, and its music video. Show me the phrases you found and the hooks first.Review the repository before installing any Skill that runs scripts. Read the source on GitHub.
Everything this makes is one kind of video: a dark stage, lights, a singer cut to the beat, the lyrics in type. Not a story told in scenes, and not footage. It is 1920×1080; it does not make a vertical one yet.
A song has an audience, not a feature list. For a repository or a product page, the launch-video Skill reads the product and narrates it.
The launch-video SkillIt is a working session with an agent, and it stops for you three times. Only you can hear the song, and only you know whether the lyric sounds like your people.
For the song and the pictures of the singer, yes. Both are made through an Angles API key and use credits from the account: 30 for a song, which comes back as two performances, and 5 for a picture. A new account starts with 200, which covers what the video on this page took to make: one song and twenty pictures. With a song file of your own and no singer, nothing needs an account.
Yes. Give it the file and the lyrics exactly as they are sung, and it starts at the measuring: bars, who is playing, when each word is sung, then the video.
Something has to hear the song, and the agent cannot. The Skill offers to install a listener on your machine — Whisper, about 3 GB, which sends nothing anywhere and has been run on Apple Silicon only. With it, every word gets the moment it is sung. Without it, you read off the player when each section starts, and each line appears whole. Either way the agent asks you to watch the moments that matter, because only you can say whether they are right.
One language a song. The two songs it was written from are in English and in Mandarin. Chinese, Japanese and Korean lyrics are timed a character at a time, everything else a word at a time.
The song is sung by a music model, and whether a made song may be used commercially is set by the terms of the service that made it. The Skill says so when it hands the video over and does not promise it for you. The lyrics are built on phrases many people share, never on one person’s sentences.
Because that is the one kind it makes. The Skill ships a stage — the lights, the way a still picture is cut to a beat, the type — and writes what they do new for every song. A different song gets different colours, cuts and a different singer on the same stage.
The video is rendered locally. The lyrics and the style of the song go to the service that sings it, and the description of each picture goes to the one that draws it. Finding the phrases reads public pages. The listener, when it is installed, runs on your machine and sends nothing.
The Skill is open source under the MIT licence. It draws the video with Remotion, which has a licence of its own — check it if you are a company. The song and the singer use credits from an Angles account; your own song and no singer use none.