Want a song? Ask for one. In an agentic chat, describing the track you
want is enough — Catalyst composes it and drops the finished result straight into the reply as
an inline audio player: press play right there in the conversation, or open the link to
save the file.
Ask for a song, a beat, a jingle, a theme, a soundtrack, or background music — or hand over
lyrics you’ve already written and ask to have them set to music. Two things shape the result:
A description — the production brief: genre, tempo, mood, instrumentation, arrangement.
Lyrics(optional) — the words to be sung. Leave them out and you get an
instrumental.
The description is a production brief, not the lyrics. Say what it should sound like:
Write a lo-fi hip-hop track, 78 BPM, D♭ major with jazzy extensions. Laid-back and dreamy, deeper in the middle, dissolving softly at the end. Warm Rhodes chords with a slow chorus wobble, dusty boom-bap drums, round sub bass, vinyl crackle. Instrumental, no vocals.
The details worth including:
Genre and style, plus tempo (BPM) and key if you care about them.
Mood, and how it evolves across the track — not just “happy,” but where it lifts and
where it settles.
Vocals: who’s singing (character, delivery, harmonies, ad-libs) — or say
“instrumental, no vocals”.
Instrumentation and production: the drums, the bass, the chords, the textures, how the
mix should feel.
Arrangement: what happens across the intro, verse, chorus, bridge, and outro.
Ask for a song with words and the assistant writes lyrics for you and sets them to music in
one step. Or bring your own:
Paste your lyrics into the chat and say what they should sound like:
Set these lyrics to music — indie folk, 92 BPM, acoustic guitar and brushed drums, a warm female lead.
Catalyst uses them verbatim. Your words aren’t rewritten; the assistant only adds
section tags if yours don’t have any.
The finished song comes back inline, with the lyrics sung over the arrangement you asked
for.
Lyrics are structured with section tags on their own lines — [Intro], [Verse], [Chorus],
[Bridge], [Instrumental], [Outro] — which tell the model where the song’s parts begin:
[Intro]
(rain on the window)
[Verse]
Six a.m. and the kettle's on
Half a thought of a half-sung song
[Chorus]
Hold the morning, hold it slow
(hold it slow)
Lines in parentheses read as backing vocals and ad-libs, and wordless Mmm… / Ooh… lines
come out as hums. Omit lyrics entirely for an instrumental — that’s the default.
Music is heavy to render, and Catalyst renders it while you wait.
Tracks are 60 seconds by default. Ask for shorter — 30 seconds is a good length for a
jingle, sting, or hook — or longer, up to the model’s cap (currently 180 seconds). The
model may also end the song naturally before the limit rather than padding it out.
Budget about five to six minutes of rendering per minute of music. A 30-second jingle
takes roughly three minutes, a 60-second track five to six, a 90-second one eight or nine.
The assistant tells you before it starts, so a long quiet pause mid-reply is normal — the
track appears as soon as it finishes.
One track renders at a time. If the shared GPU is busy with someone else’s render, the
assistant says so instead of silently queueing — just ask again in a few minutes.
You don’t normally pick a model — the assistant uses the default. Workspace admins curate which
music models are available under Admin → Music Models.
MiniMax Music 3 — songs and instrumentals (default)
The default, and today the only curated model. Full songs with vocals and lyrics, or
instrumentals; a wide genre range; and it responds well to detailed production briefs —
the more specific your description, the more it rewards you. Name it — “use MiniMax
Music” — if you ever need to be explicit.
Be concrete about sound. Name the instruments, the tempo, the texture. “Dusty boom-bap
drums and a round sub bass” beats “chill music” every time.
Say how the track moves. Where it builds, where it drops back, how it ends — an
arrangement makes it a song rather than a loop.
Keep lyrics singable. Short lines, natural rhythm, a chorus that repeats — and sized to
the length you asked for.
Ask for variations, not a batch. One request makes one track. If you want options, ask
for another take with one thing changed (“same track, but brighter and 10 BPM faster”)
instead of expecting three at once.
Write lyrics in the language you want sung — the description itself works best in English.
English and Mandarin Chinese lyrics are proven; other languages are worth a try.
Any genre, and duets. From trap and festival house to orchestral trailer cues, jazz,
country, metal, Mandopop and gospel — name the genre (not just a mood) and the assistant
writes the brief in the model’s own idiom. For a duet, say who sings what (“male verses,
female chorus, both on the last chorus”).
A finished track renders as a native audio player inline in the reply, with standard
controls, and it travels with the conversation — anyone viewing a
shared chat can play it too. The file itself is an .mp3 on Catalyst’s
media CDN; open the link to save or share it.
Under the hood, plain language is all you need: a music-authorskill loads automatically and teaches the assistant how to write the
production brief, format lyrics with section tags, and size the track to what you asked for —
so you can just ask for the song you want.