← All guides
API for GenAI

The AI Music Video Prompt: How Songbrain Powers GenAI Video Tools

May 14, 2026 · Updated October 2026 · 8 min read

Second-level sync: song event in, shot change out

The drops at 0:22 and 1:24, the best moment at 0:22–0:42 and the breakdown at 0:42 are the exact seconds the scene cuts land onSONG · the analyzed energy curve“Night Lap” · drift phonk · 2:46★ DROP · 0:22★ DROP 2 · 1:24VIDEO · one scene block per eventS1S2S5S6S7S9S10★ BEST MOMENT · 0:22–0:42★ OUTRO PEAK · 1:58–2:460:000:220:421:241:582:46

The dashed lines are the whole product. A video model has no reason to cut on 0:22 — it can't hear the drop. The spec below hands it that second, and the nine others, before it renders a frame.

From song analysis to rendered music video

An AI music video prompt is only as good as its knowledge of the song: the sec-precise analysis comes first, the GenAI tool renders last.

Key takeaways

  • →Video-gen models are no longer the bottleneck — the prompt is, and a music-video prompt has to know the song second by second.
  • →Songbrain hands the video tool best moments, energy curve, mood signature, subgenre aesthetics, target audience and lyric hooks in one API call.
  • →The scene spec anchors every cut to a timestamp — the drop gets the visual pop, the choruses get a deliberate visual rhyme, the breakdown gets stillness.
  • →Clip export targets for TikTok, Reels and Shorts are pre-computed from the same analysis, so no human has to pick the clip.

Sora 2 renders cinematic 4K. Runway Gen-4 is up to 20 seconds per shot. Veo 3 nails physics. Kling 2 shoots clean dolly-ins. The video generation models are no longer the bottleneck. The prompt is.

And when the goal is a music video, the prompt has a problem no general-purpose video tool can solve on its own: it needs to know the song. Sec-precise. Where the drop is. Which moments are clip-worthy. What the dominant mood is at second 22 vs. second 88. Who's actually going to watch it.

That's the layer Songbrain owns. Our long-term goal is to be the best music-intelligence API for genAI tooling — specifically the one video-gen tools call before they render a frame. This guide walks through what we hand back, and what a Sora/Runway/Veo/Kling-class model does with it.

What a video-gen AI gets from one API call

best_momentsSec-precise start/end of each hook, with confidence score and moment type (intro_grab, pre_drop, chorus, drop, outro_peak).
energy_curveTime-series energy 0–1 across the full song, plus the exact moments where it inflects.
mood_signatureDominant moods per section (aggressive, euphoric, nocturnal, melancholic, hype, tense). The emotional script for the visual.
subgenre + style_descriptorsVisual aesthetic genome — e.g. drift phonk → chrome / sodium-vapor / JDM cars / dust particles / 90s VHS overlay.
target_audienceDemographics, regions, platforms and feed contexts derived from mood, subgenre and the song's DNA.
lyric_hooksThe strongest repeating lyric lines, with timestamps. The lines the visual should literalize or contradict for maximum tension.
genre_styleVisual style cues typical for the subgenre, so the AI leans into feed-fit instead of a generic look.

From data to scene script

On its own, that data dump is just a JSON blob. The magic happens in the next step: Songbrain compiles it into a scene-by-scene video prompt that speaks the language of modern video-gen models. Sec-precise. Camera-aware. Aesthetic-locked. Audience-tuned.

Below is a full example. The song: a fictional drift phonk track called “Night Lap,” 2:46, Virality Score 88. This is the exact prompt-spec a video-gen tool would receive when calling Songbrain's API and asking “turn this into a music video.”

Example: “Night Lap” — Songbrain video prompt spec

The 20 seconds the whole spec is built around

★ BEST MOMENT · 0:22–0:42

Scene 5 in the spec below — the primary chorus, energy 0.92, flagged as “the segment 90% of the clips will sample from.” Every other scene is written around it: the pre-drop withholds, the breakdown recovers, the second chorus rhymes with it.

GET /api/v1/genai/video-spec/{job_id}
{
"song_id": "sb_a3d9_night_lap",
"duration_sec": 166,
"subgenre": "Drift Phonk",
"virality_score": 88,
"primary_moods": ["aggressive", "nocturnal", "tense", "hype"],
"global_aesthetic": "sodium-vapor orange + neon cyan, crushed blacks, JDM coupe culture, 21:9 letterbox, 90s VHS chroma bleed, dust particles in cone-of-light, light leaks on every drop",
"target_audience": {
"demo": "male-leaning, 16–26",
"regions": ["Eastern Europe", "LATAM", "SEA", "US car-culture pockets"],
"platforms": ["TikTok", "YouTube Shorts", "Instagram Reels"],
"feed_contexts": ["drift edits", "anime AMV", "gym hype", "JDM car compilations"],
"trend_tags": ["#driftphonk", "#nightlap", "#jdm", "#phonkedit"]
},
"scenes": [
{
"t": "0:00 – 0:08",
"energy": 0.35,
"mood": "tense",
"prompt": "Slow dolly-in toward a black-tinted Nissan S15 parked in a wet underground garage. Sodium-vapor orange overhead, neon cyan reflections on the hood. Engine off. Steam rising from the asphalt. Cinematic 21:9. 35mm anamorphic lens, shallow depth of field. No subject yet — just the car."
},
{
"t": "0:08 – 0:18",
"energy": 0.55,
"mood": "tense rising",
"prompt": "Quick cuts (0.8s avg): male hands on steering wheel (rings, scarred knuckles), key turning in ignition, hand on gear shift, lighter clicking. Bleach-bypass color grade, crushed blacks. Each cut on the cowbell hit. No face visible yet — withhold the subject."
},
{
"t": "0:18 – 0:22",
"energy": 0.85,
"label": "PRE-DROP BUILD",
"mood": "tension max",
"prompt": "Engine starts. Camera pushes in 24fps slow on the driver's eyes only — male, late 20s, sharp jaw, dead-serious into the lens. Eyes lit from below by the dashboard. Dust particles drifting through the cone of light. Hold the stare. The world goes quiet for half a beat."
},
{
"t": "0:22",
"energy": 1.0,
"label": "★ DROP — visual pop",
"mood": "aggressive release",
"prompt": "HARD CUT TO: tires breaking traction in slow-mo 240fps. Smoke geyser from the rear wheels. Reverse-shot from a low front-bumper rig as the car launches forward. RGB chromatic aberration on each rim spoke. Single hard lens flare across the windshield. Strobe-frame flash of the headlight cutting on. Letterbox cracks open one frame wider on this beat. This is the most clip-worthy frame in the whole video — designed for the TikTok freeze-frame."
},
{
"t": "0:22 – 0:42",
"energy": 0.92,
"label": "★ BEST MOMENT — primary chorus",
"mood": "aggressive + euphoric",
"prompt": "Drift sequence on a multi-level neon parking deck. Whip-pan follow shot, then mounted side-cam through the door window — driver's profile lit by the strip lights flickering past. Tire-smoke catches the cyan and goes magenta where it hits the orange. Cut on every 4th beat: wide drift / driver close-up / wheel slip / sparks. Letterbox stays. Color grade: peak saturation, no blooms — keep it punchy, not dreamy. This is the segment 90% of the clips will sample from."
},

The same render, pre-cut for three feeds

Clip export targets are pre-computed from the same analysis — no human picks the clipTikTok primaryReels secondaryShorts outro tease24s15s14s0:18 – 0:421:24 – 1:392:32 – 2:460:002:46No human picks the clip.

These three windows come out of editor_notes.clip_export_targetsat the end of the same spec — computed from the Best Moments, not chosen by an editor scrubbing a timeline.

…/video-spec/{job_id}— scenes 6–10 + editor notes
{
"t": "0:42 – 1:08",
"energy": 0.45,
"mood": "nocturnal, hollow",
"prompt": "Cut to static wide on a rain-slick bridge overpass. Single sodium light. Car parked. Driver stands at the rail, smoking, back to camera, looking down at the city. No movement for 6 full seconds — let the breakdown breathe. Then slow handheld push-in on his shoulders. VHS tracking artifacts at the edges of the frame. Mood: he's alone with whatever this song is about."
},
{
"t": "1:08 – 1:24",
"energy": 0.78,
"mood": "building anger",
"prompt": "Back in the driver seat. Quick intercuts: rearview mirror (red and blue flashing far behind, just a hint), foot pressing the pedal harder, shifter slamming into gear, dashboard tach climbing to redline. Cut tempo accelerates from 1.2s shots down to 0.4s shots over the 16-second window."
},
{
"t": "1:24",
"energy": 1.0,
"label": "★ DROP 2 — chase release",
"mood": "high-octane",
"prompt": "Same visual pop as scene 4 — but now the car is moving. Top-down drone tracking shot of the S15 weaving between cones / barriers / sleeping traffic. Cut to: side-mounted shot, full slide, smoke wall obscures everything for half a second, smoke clears on the next beat, car still in frame."
},
{
"t": "1:24 – 1:58",
"energy": 0.95,
"mood": "aggressive + euphoric (callback)",
"prompt": "Variation on scene 5's chorus visual but at street-level: long empty boulevard, palm-tree silhouettes against orange sky, drift-line cutting across the centerline. Reuse the side-cam-through-window angle on the driver — same shot, different city. The callback is the point: the audience should recognize the visual rhyme between the two choruses."
},
{
"t": "1:58 – 2:46",
"energy": 0.5,
"label": "★ OUTRO PEAK — emotional close",
"mood": "spent, melancholic, defiant",
"prompt": "Sunrise pulling the orange out of the sky. Car parked on an empty cliff road overlooking the city. Engine off, ticking from heat. Driver's silhouette against the dawn — same eyes from scene 3 now reflective, not aggressive. Last 6 seconds: slow dolly out, the car gets smaller, the city wakes up, the music fades on a held cowbell. Hard cut to black."
}
],
"editor_notes": {
"never_show": ["the antagonist", "lyrics on screen", "stock dashcam stock-footage"],
"keep_letterbox_throughout": true,
"clip_export_targets": [
"0:18 – 0:42 (TikTok primary, 24s)",
"1:24 – 1:39 (Reels secondary, 15s)",
"2:32 – 2:46 (Shorts outro tease, 14s)"
]
}
}

Why this matters for video-gen tools

A video-gen model on its own has no reason to cut on beat 22. It doesn't know the chorus comes back at 1:24. It doesn't know that drift phonk audiences expect the visual rhyme between chorus one and chorus two, or that the breakdown demands stillness. It guesses. The result — even from Sora 2 or Veo 3 — is technically clean but emotionally arbitrary.

Songbrain's job is to remove the guessing. Every video-gen company building a “music video” product is going to need this layer. The choice is: build the music intelligence stack themselves (a multi-stage audio, lyrics and moment pipeline), or call our API.

Same renderer, one API call apart

Nothing about the renderer changes. The only thing that changes is whether it knows the song — which is the difference between footage and a music video.

What else changes when you add Songbrain

  • Clip export targets are pre-computed. The same render is auto-trimmed into TikTok, Reels and Shorts cuts using Songbrain's Best Moments timestamps. No human picks the clip.
  • Audience-aware visual choices. A K-pop video gets a different aesthetic library than a deathcore one — not because the model picked it, but because Songbrain's subgenre + audience layer told it to.
  • Genre-aware aesthetics. Drift phonk leans into VHS chroma bleed and night drives, not glossy CGI — the prompt carries that from the subgenre instead of leaving it to chance.
  • Caption + hashtag handoff. The same API call also returns the social-media kit: caption, hashtags, hook line. The video and the post share one source of truth.

Who this is for

Three kinds of customer, all building the same gap:

  • Video-gen platforms adding a “music video” mode (Sora, Runway, Veo, Kling-class tools wanting better audio-aware output).
  • AI music video startups building consumer products on top of Sora/Runway/Veo APIs.
  • Label / DSP / promo tools wanting auto-generated promo videos per release without hiring an editor for every track.

We're onboarding integration partners now. The API returns the JSON above — plus the original full analysis (Virality Score, Best Moments, Song DNA, lyrics, genre) — in about a minute per song. Just want the finished video for your own song? Songbrain's one-click AI music video does this end to end.

Building a music-video-gen tool?

Call the Songbrain API once per song: beat grid, best moments, timed lyrics and a beat-synced shot plan with a prompt per scene. 5 free songs a month.

Get a free API key →

Find the Viral Radar for your genre

The best Songbrain songs in your genre, and where your own track would land.

Free · no account · ~60 seconds

Get the prompts for your own song

The video prompts are generated from your track's own moments, tempo and mood — upload it and they come back with the analysis.

Drop your song here
MP3 or WAV · 30s minimum · free, no account
0–100 scorebest moments7 reelsposting plan~60s

Tools for this