angles.video

create-music-video · open-source Agent Skill

Concert-stage music videos your coding agent writes, has sung, and cuts to the song

I Will Not Promote

1:29 · Sung in English · the hook is a line 344 different people had typed

It makes one kind of music video: a dark stage, lights that follow the band, every lyric in type as it is sung, and a drawn singer if you want one. Not generated footage, and not a story told in scenes.

Name an audience. The agent reads where they talk, builds the lyrics on phrases they really say, and has a music model sing them. Then it measures the song — its bars, who is playing, when each word is sung — and writes the stage against those measurements in Remotion. It renders on your machine.

Real output

How the one above was made

This is the video the Skill was written from. We named an audience — solo founders who build and never promote. The first lyric the agent wrote, without reading anything, we turned down. So it read 1,849 posts from nine founder communities and counted: the hook is a line 344 different people had typed. We agreed the lyrics, chose the second of two performances, turned down four singers for being too good-looking, and picked one of the next four. Then the steps were written down as the Skill.

Every video the Skill makes is this kind. Another song gets other colours, other cuts and another singer, on the same stage.

One session, start to finish
  1. youMake a song for solo founders who build and never promote, and its music video.
  2. agentRead 1,849 posts from nine founder communities and counted. The phrase the most people typed: “I will not promote” — 344 people.
  3. agentThe lyrics: 22 lines, with the phrase under each line and how many people say it.You agree
  4. agentTwo performances of it are in public/song/. Which one?You listen
  5. agentMeasured: 165 beats a minute, 60 bars. The band stops in bars 14 and 40 — the hook goes there, alone.
  6. agentFour singers to choose from, side by side.You pick
  7. agentout/promote.final.mp4 — 1:29, 1920×1080.

Numbers from the video above: I Will Not Promote

It stops for you three times: the words, the song, the singer

  1. 01

    You name an audience

    “A song about my app” is a topic. “A song for solo founders who never promote what they build” has an audience, and the lyrics are going to come out of their mouths. Bring a song file of your own instead, and it starts at the measuring.

  2. 02

    It reads where they talk, and counts

    How many different people say a phrase in the same words, how recently, and what it means where it is said. One clever post is one person’s writing. A line three hundred people typed is a chorus.

  3. 03

    You agree the lyrics

    Two or three hooks first, each with its count. Then the lyrics with the phrase under every line — before any song is made, because a lyric you do not like is not rescued by a good tune.

  4. 04

    You choose the performance

    A music model sings the words as written, twice. The agent cannot hear either one. You listen to both, pick, and say whether anything is shouted that is not in the lyrics.

  5. 05

    You pick the singer, or have none

    Three or four candidates that really differ, drawn and shown together, because nobody can choose a face from a description. Then every shot is drawn from the one you picked.

  6. 06

    It measures, writes, looks, and fixes

    Bars, who is playing in each, when each word is sung. The lights and the cuts are written against that. It renders the frames that have to be right, looks at them, and fixes what is wrong before you see it.

Nothing is laid over the song by guesswork

The lights follow who is playing
A dense song is equally loud from its fifth second on. What changes is who plays: no kick under the first verse, a kick once a bar, then on every beat. Each look in the video starts where one of those changes does.
Where the band stops, the lights stop
Black, and the largest type in the video, one word at a time as it is sung. Then everything at once when the band comes back. If the song has no such bar, the video does not invent one.
A word appears when it is sung
Its time is heard, not laid on a grid. Where the voice is alone, it is taken from the sound itself, because a listener starts every shouted word up to a fifth of a second early.
You are told what was not heard
Every line is reported with how its time was come by. A line that was not found has no time at all, and you are told which moments to watch with the sound on.
Never more than three flashes a second
Faster than that harms viewers with photosensitive epilepsy. Fast rhythm is carried by lights that move — a chase, a sweep — not by the whole frame blinking.
// src/promote.phrases.json
{
  "phrase": "I will not promote",
  "people": 344,
  "newest": "2026-09-16",
  "where": "r/startups",
  "means": "the line every post there has to carry in its title; people resent typing it"
}
 bar    from   kick                                voice and chords
  13  18.26s   ##       ##       ##       ##          #####    #####    #####    ####
  14  19.72s   #                                      ######   ######   ######   #####
  15  21.20s   ####     ###      ###      ##          ####     ####     #####    #####
// src/promote.lines.json
{
  "section": "chorus 1",
  "text": "I will not promote",
  "how": "heard",
  "by": "its own sound",
  "seconds": [19.67, 20.11, 20.47, 20.83]
}
From the song above: the hook in its phrases file, three bars of what the measuring printed — the kick is gone from bar 14 — and that hook as it was timed.

What is on the stage

Moving heads

A row above the frame and a row on the floor. Where each one points and how bright it is, is written for this song, bar by bar.

A room with air in it

Colour thrown across the walls and floor, and haze drifting through it — which is what makes a beam visible at all.

A singer made of still pictures

One drawn person in eight to twelve shots. What makes them move is the cutting: each picture framed several ways, each cut landing on a bar or a word.

Lyrics that arrive as they are sung

Told lines fill in along the bottom, word by word. Shouted lines are set as large as the frame will hold, and land a word at a time.

Things held back

Sparks, paper, lasers, and one full-frame white flash. The second chorus has to have somewhere to go that the first did not.

A readout

A strip of meters moving to the song’s own spectrum, and a frame of readings that says what the lyric says in another voice.

How the words are timed

With a listener on your machine
Whisper, about 3 GB, installed into the workspace only if you agree to it. The song is not sent anywhere. Every word gets the moment it is sung. It has been run on Apple Silicon only.
Without one
You read off the player when each section starts. Each line then appears whole, at about the right bar — and nothing arrives a word at a time.
What neither can do
Tell whether the words are in time with the voice. The agent has times and no ears, so it names the moments that matter and asks you to watch them.

What you are handed

promote.final.mp4
The video, at publishing loudness.
promote.lyrics.txt
The words, as sung.
promote.phrases.json
The phrase under each line, and how many people say it.
promote.lines.json
When each word is sung, and how that was come by.
promote.tsx
The lights and the cuts, yours to change.
public/song/
Both performances of the song.

Install it

  • Codex or Claude Code
  • Node.js 18 or newer
  • Python 3 with numpy, and ffmpeg
  • An Angles API key for the song and the singer — or a song of your own

Codex

$skill-installer install https://github.com/anglesvideo/angles-video-skills/tree/main/skills/create-music-video

Claude Code

git clone --depth 1 https://github.com/anglesvideo/angles-video-skills.git ~/angles-video-skills
mkdir -p ~/.claude/skills
ln -s ~/angles-video-skills/skills/create-music-video ~/.claude/skills/create-music-video

Then ask

Make a ninety-second song for people who run a one-person newsletter, and its music video. Show me the phrases you found and the hooks first.

Review the repository before installing any Skill that runs scripts. Read the source on GitHub.

When it is not the right tool

A video that is not a performance

Everything this makes is one kind of video: a dark stage, lights, a singer cut to the beat, the lyrics in type. Not a story told in scenes, and not footage. It is 1920×1080; it does not make a vertical one yet.

A video about your product

A song has an audience, not a feature list. For a repository or a product page, the launch-video Skill reads the product and narrates it.

The launch-video Skill

One click and done

It is a working session with an agent, and it stops for you three times. Only you can hear the song, and only you know whether the lyric sounds like your people.

Questions about the music video Skill

Do I need an Angles account?

For the song and the pictures of the singer, yes. Both are made through an Angles API key and use credits from the account: 30 for a song, which comes back as two performances, and 5 for a picture. A new account starts with 200, which covers what the video on this page took to make: one song and twenty pictures. With a song file of your own and no singer, nothing needs an account.

Can I use a song I already have?

Yes. Give it the file and the lyrics exactly as they are sung, and it starts at the measuring: bars, who is playing, when each word is sung, then the video.

How do the words end up in time with the voice?

Something has to hear the song, and the agent cannot. The Skill offers to install a listener on your machine — Whisper, about 3 GB, which sends nothing anywhere and has been run on Apple Silicon only. With it, every word gets the moment it is sung. Without it, you read off the player when each section starts, and each line appears whole. Either way the agent asks you to watch the moments that matter, because only you can say whether they are right.

Which languages can the song be in?

One language a song. The two songs it was written from are in English and in Mandarin. Chinese, Japanese and Korean lyrics are timed a character at a time, everything else a word at a time.

Can I release the song?

The song is sung by a music model, and whether a made song may be used commercially is set by the terms of the service that made it. The Skill says so when it hands the video over and does not promise it for you. The lyrics are built on phrases many people share, never on one person’s sentences.

Why does every video look like a concert?

Because that is the one kind it makes. The Skill ships a stage — the lights, the way a still picture is cut to a beat, the type — and writes what they do new for every song. A different song gets different colours, cuts and a different singer on the same stage.

Does anything leave my machine?

The video is rendered locally. The lyrics and the style of the song go to the service that sings it, and the description of each picture goes to the one that draws it. Finding the phrases reads public pages. The listener, when it is installed, runs on your machine and sends nothing.

What does it cost?

The Skill is open source under the MIT licence. It draws the video with Remotion, which has a licence of its own — check it if you are a company. The song and the singer use credits from an Angles account; your own song and no singer use none.