Hotel Lobby AI Prompt: The Exact Prompt Our Generator Uses
Most Hotel Lobby AI prompts online are guesses. These are not: they are the texts our generator sends for Hotel Lobby videos, printed from the same code. The first is for our default model, MiniMax H3 reference-to-video; the second is the Seedance version that writes an original rap. Copy them, adapt them to your tool, or skip them and let the generator do the work.
Updated
The prompt
This is the version for a vertical 12-second video with a man on the left and a woman on the right. The model receives four attachments in this order: Image 1 is the left photo, Image 2 the right photo, Video 1 the Hotel Lobby reference clip and Audio 1 the reference track.
Recreate @Video1 shot for shot: the same orange gradient backdrop with the round glowing halo behind the performers, the same single silver condenser microphone hanging from the ceiling between them, the same camera framing, timing and every movement, 9:16.
Performer on the left: the man from @Image1. Performer on the right: the woman from @Image2. Keep each person's face, hairstyle and outfit exactly as in their photo. Use only the people from the photos; ignore the photo backgrounds completely.
The left performer raps into the hanging microphone for the whole clip, lip-syncing precisely to @Audio1, with natural head nods and expressive mouth movement. The right performer is the hype partner: she points at him, reacts, smiles and makes hand gestures on the beat, mouthing only short ad-libs. Around the 7-second mark the camera pulls back to a wider waist-up shot and she does a short playful dance move, as in @Video1.
Single continuous shot. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.The AI rap version (Seedance)
On a Seedance model with the AI rap sound, no track is attached: the model writes the song itself. The reference clip goes up without its sound, so the model copies the moves and the booth but not the Hotel Lobby recording. Attachments are named in words (“reference image 1”), because kie documents no mention syntax for Seedance. This one is for the example topic below.
Recreate reference video 1 shot for shot: the same orange gradient backdrop with the round glowing halo behind the performers, the same single silver condenser microphone hanging from the ceiling between them, the same camera framing, timing and every movement, 9:16.
Performer on the left: the man from reference image 1. Performer on the right: the woman from reference image 2. Keep each person's face, hairstyle and outfit exactly as in their photo. Use only the people from the photos; ignore the photo backgrounds completely.
Soundtrack: an original hip-hop song made for this video, a punchy trap beat with deep bass and crisp hi-hats. The left performer raps an original, clean verse into the hanging microphone about: "Maya’s 30th birthday and her terrible parallel parking". Rap in the language of that topic. Clear rap vocals, lips in sync with every word, natural head nods and expressive mouth movement. The right performer is the hype partner: she points at the rapper, reacts, smiles and adds short ad-libs on the beat. Do not use any existing song, existing lyrics or the voice of any real artist. Around the 7-second mark the camera pulls back to a wider waist-up shot and she does a short playful dance move, as in reference video 1.
Single continuous shot. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.What each part does
- The scene. “Recreate @Video1 shot for shot” hands the booth, the camera and the timing to the reference clip, and the prompt names the parts that matter most: the orange gradient backdrop, the round glowing halo and one silver condenser microphone hanging between the performers.
- The identities. Each photo is tied to one side, and the model is told to keep face, hairstyle and outfit and to ignore the photo backgrounds. Without this line, performers swap sides or wear the reference clip’s clothes.
- The roles. The left performer raps into the mic, lip-syncing to the audio; the right one is the hype partner who points, reacts and mouths only short ad-libs. Clips of 8 seconds or more add the camera pull-back and a short dance near the 7-second mark, as in the reference.
- The guard rails. One continuous shot, no split screen, no extra people, no captions, faces consistent from start to end. These stop the most common failures.
How to use it in another AI video tool
- It needs a model that accepts reference images and a reference video, and ideally a reference audio track. MiniMax H3 reference-to-video and ByteDance Seedance 2.x take all three.
- Change the mentions to your tool’s syntax. We write @Image1 and @Video1; some tools expect “Image 1” and “Video 1”, others their own tags. Keep the order: left photo first, right photo second.
- Change the performer nouns. Our generator swaps only “man” and “woman” and the matching he or she; every other word stays fixed so results are comparable.
- Change the shape in the first line (9:16, 16:9, 1:1) to match the format you pick in the tool.
- Without a reference clip, describe the booth and the moves in words instead of “Recreate @Video1”. Tools such as Dreamina publish their own text-only prompts for this. We have only tested the version above.
Common problems and fixes
| Problem | Likely cause | Fix |
|---|---|---|
| Faces drift or look like someone else | Side profile, sunglasses or a small face in the photo | Use front-facing, well-lit photos, waist up |
| The two people swap sides | The prompt does not tie each photo to a side | Name the left and right performer with their image |
| A third person appears | The model copies people from the reference or a photo background | Keep “no extra people” and “ignore the photo backgrounds” |
| Words appear on screen | The model adds captions | Keep “no captions or text on screen” |
| Lips do not match the audio | No audio attached, or the audio starts on silence | Attach the track and start it on a word |
Or skip the prompt
On Hotel Lobby Video this prompt, the reference clip and the reference track are built in. You add two photos, pick He or She for each spot, and create. The first preview is free: one HD video of up to 12 seconds, with a watermark, once per Google account.
FAQ
What is the best prompt for the Hotel Lobby AI trend?
One that fixes the booth, the single hanging mic, which photo stands on which side, and what each person does. The prompt above is the one our generator uses for every video.
Which AI model does this prompt work with?
We use the first with MiniMax H3 reference-to-video and the AI rap version with Seedance 2.x; both take reference images, a reference video and a reference audio track. Other models need their own mention syntax and may not accept all three.
Do I need a prompt to make a Hotel Lobby AI video?
Not on a two-photo generator. On Hotel Lobby Video the prompt is built in; you only add the photos.
Sources
Read next
How to Do the Hotel Lobby Trend: 3 Ways, Step by Step (2026)
Three ways to do the Hotel Lobby AI trend (two-photo generator, AI video prompt, editor template), with steps, costs, free options and tips
Best Hotel Lobby AI Video Generators (Checked Sep 30, 2026)
A dated comparison of Hotel Lobby AI video generators (Hotel Lobby Video, Rap Duo AI, LightX, Dreamina, Media.io, Summrs, Fuzana, Kapwing) from their own pages
Hotel Lobby Video is an independent tool, not affiliated with Quavo, Takeoff, Migos, Quality Control Music, Motown or COLORS.