8 組真的有效的 Grok Imagine 提示詞
八組可以直接貼上的 Grok Imagine 1.5 提示詞,圖片和影片都有——建立在 FLUX.1 的自然語言風格上,用主體 + 動作 + 運鏡 + 風格 + 音訊的結構。每一組做什麼、為什麼成立,都在 OmniArt 上跑得動。

Grok Imagine 1.5 把圖片底模換成了 Black Forest Labs 的 FLUX.1,這個改變對你怎麼寫提示詞有很具體的影響:這個模型對自然語言描述的反應,比較像攝影師讀一份簡報,而不是舊模型那種解析關鍵字清單的方式。下面八組提示詞可以直接貼上——丟進 OmniArt 的 Grok Imagine 工作台,把細節改成你的,然後生成。每一張卡片都附上完整的提示詞原文、它會產出什麼,以及一則關於這個結構為什麼成立的手藝筆記。
想看橫跨所有 OmniArt 模型的通用提示詞理論,請看怎麼寫出更好的提示詞。想深入了解 Grok Imagine 的六種生成模式和成本計算,請看 Grok Imagine 創作者指南。這篇文章講的是 Grok Imagine 1.5——也就是 FLUX.1 那一版——以及它會回報你的那種提示詞手藝。
Grok Imagine 1.5 改變了什麼寫法
FLUX.1 這個底模的訓練方式和早期的文生圖架構不一樣。它很會解析連貫的散文,對純關鍵字堆疊反而反應不足。有五個習慣最能穩定拉高品質:
- 用自然語言,不要堆關鍵字。 完整句子勝過用逗號串起來的形容詞。"A street at blue hour, lit by the hum of a convenience store sign" 打得贏 "street, night, neon, cinematic, 4K."
- 用具體參照,不要模糊形容詞。 "Shot on a Fujifilm XT4, 23mm f/2" 告訴模型的資訊,比 "high quality photo" 多得多。點名的器材和底片,在潛空間裡是有真實重量的。
- 用精確的顏色詞,不要說「colorful」。 "Electric blue and hot pink" 產出一組刻意選過的色盤。"Colorful" 產出的是被平均掉的雜訊。
- 用精確的時間,不要說「golden hour」。 "Late October, 5:45 pm, sun 6° above the horizon" 告訴模型光線的確切角度和色溫。"Golden hour" 跨季節、跨緯度都是含糊的。
- 影片的結構:主體 + 動作 + 運鏡 + 風格 + 音訊。 把核心的主體和動作放進前 20–30 個字。單一風格焦點勝過混搭。逐步迭代——每次生成只改一個變數,直到結果穩住,再往前推。
想完整了解那套能轉移到影片上的電影語彙,電影感 AI 影片提示詞指南深入談了鏡頭選擇、有動機的運鏡和光線用語。
八組提示詞
1. 電影感商品照(圖片)
35mm product photography, shot on Fujifilm XT4. A matte black mechanical wristwatch resting on a slab of raw concrete,
late October afternoon light coming in low from camera left at roughly 20°, casting a long shadow across the concrete
face. Shallow depth of field, background falling completely soft. Color palette: warm amber highlights, cool blue-grey
shadow fill. No props, no reflections except the concrete surface itself.
它會產出什麼: 一張乾淨、有藝術指導感的靜態圖,讀起來像專業商品攝影,而不是 AI 的輸出。
為什麼成立: Fujifilm XT4 這個參照把色彩科學和感光元件的成像,錨定在一種真實世界的具體樣貌上。光線角度用數字指定,避免模型退回預設的頂部散射光。把色盤限制在兩個顏色——暖琥珀色的高光、冷藍灰的暗部補光——可以阻止模型自己加進第三個互相打架的色相。
2. 帶音訊的角色特寫(影片)
Medium close-up of a young woman with short silver hair and a worn leather jacket, inside a neon-lit record shop at
3 am. She looks directly into camera and says: "Every city has one song. I'm still looking for mine." Natural lip
sync. Camera holds completely still. Light source: one pink neon tube overhead, one cyan neon sign spilling from
camera right. Atmosphere: quiet, a little melancholic, not cinematic drama. Ambient audio: low vinyl static underneath
the dialogue. 8 seconds.
它會產出什麼: 一個帶 Grok Imagine 1.5 原生音訊的角色時刻——模型在同一次推論裡生成對白、對嘴和環境音。
為什麼成立: 這句對白夠短,在 8 秒內可以乾淨對嘴。兩個分開命名的霓虹光源(頭頂粉紅、右側青色)給了模型清楚的光線地圖,避免它平均成一團通用的「霓虹都市」。"Not cinematic drama" 是一個否定式限制,它引導氛圍的精準度,比任何肯定式形容詞都高。
小技巧
10 秒以內的片子,對白請控制在一到兩個短句。句子太長會把可用時長塞爆,模型可能會把台詞唸得很趕,或是把音訊提早切掉。
3. 氛圍環境——空景片(影片)
Wide establishing shot of a fog-filled pine forest in southern Norway, early November, 7 am. No people, no animals.
Soft diffused dawn light filtering through the canopy, pale grey-white, casting almost no shadow. Slow imperceptible
push forward, as if the camera is drifting on breath. Audio: deep forest ambience — distant water, occasional bird,
near-silence underneath. No music. 12 seconds.
它會產出什麼: 一段定調氣氛的空景,很適合當背景畫面、轉場素材或開場鏡頭。
為什麼成立: "early November, 7 am" 比「霧氣的早晨」精準。推鏡被形容成 "imperceptible" 和 "drifting on breath",這比「慢慢推近」更準確地傳達了節奏。要求不要音樂,可以避免音訊退回預設的配樂——模型會改成生成真正像實地收音的環境音。
4. 快節奏直式社群——商品揭曉(影片)
9:16 vertical. A pair of electric blue running shoes drops into frame from the top, landing on a wet reflective black
studio floor. High-speed impact, tiny water spray, shoes bounce once and settle. Immediate cut to product floating
at centre frame, slow rotation 360°. Fast rhythm: first motion 0–2s, rotation 2–8s. Hard direct light from above,
electric blue accent light from below floor (subtle). No dialogue. Audio: sharp impact sound on drop, then a clean
single synthesizer tone during rotation. 8 seconds.
它會產出什麼: 一支有力道的 9:16 社群短片,是為 TikTok、Reels 或 Shorts 做的——快剪的商品揭曉,帶原生音訊。
為什麼成立: 把 9:16 放在最前面,等於在提示詞的其他東西之前先定下畫面比例。時間軸被明確寫出來("0–2s / 2–8s"),這幫助模型把兩個節拍分開拿捏,而不是把它們混成一個動作。點名具體的音訊事件(撞擊聲、合成器音)產生的音效設計,比「加點音效」有意識得多。
注意
Grok Imagine 1.5 的片子最長 15 秒。做社群內容的話,請把片長控制在 8–10 秒以內——模型的動態在這個區間最乾淨,社群平台的注意力窗口也很短。在 720p 下,一支 8 秒的片子在 OmniArt 上要 120 點數。
5. 風格化插畫(圖片)
Risograph print illustration of a small coastal Japanese fishing village at dusk, mid-December. Two ink colors only:
deep indigo and warm persimmon orange. Flat graphic shapes, no gradients. Fishing boats pulled up on shore, a single
wooden dock, lantern light in two window rectangles. Composition: low horizon line, large sky area, boats and dock in
lower third. The print has slight ink misregistration — indigo shifted 2px left from the orange layer. Texture:
visible paper grain throughout.
它會產出什麼: 一張圖像感強、限定用色的插畫,讀起來像真的印刷製程,而不是通用的數位繪圖。
為什麼成立: 點名印刷工藝(Risograph)以及它的具體限制(兩色油墨、平塗色塊、無漸層、套印偏移),等於給了模型一份完整的技術簡報。"Ink misregistration" 就是那種能把輸出錨在真實世界美感上的物理製程細節——它在 FLUX.1 上的地位,等同於點名一支底片。少了它,模型就會傾向加上漸層或把顏色混在一起。
6. 動態運鏡——空拍後拉(影片)
Aerial drone footage. Extreme close-up on the face of a compass resting on a weathered wooden ship's deck, late
afternoon November light, warm golden horizontal rays from camera left. Slow pull-back revealing the full deck,
then the ship's hull, then open grey Atlantic ocean horizon. Pull-back runs the full 15 seconds — begin on compass,
end with ocean filling 80% of the frame. Camera elevation stays constant, no tilt. Real drone color science: flat
LOG-style color, slight lens vignette. Audio: wind increasing in volume as ocean fills frame.
它會產出什麼: 一顆撐滿 15 秒的揭曉鏡頭——也就是這個模型的最長片長——整支片圍繞著一個有動機的運鏡。
為什麼成立: 這組提示詞把完整的 15 秒用在一個連續動作上,這是在這個長度拿到乾淨結果最可靠的方式。後拉被限制在固定高度(不做俯仰),避免模型自己即興加上第二個運鏡軸,把動態弄得卡卡的。"LOG-style color, slight lens vignette" 編碼出真實攝影機的樣貌,又不用點名具體器材。
7. 風格化時尚——底片肖像(圖片)
Expired Kodak Portra 400 film scan. Portrait of a woman in her mid-thirties, strong afternoon window light from
camera right, half of her face in deep shadow. She is wearing a deep forest green linen blazer, no visible jewellery.
Expression is neutral, looking slightly off-camera left. Grain heavy and warm, slight halation around the window
highlight, greens shifted slightly toward yellow-olive. Tight crop: from collarbone to just above top of head.
Aspect ratio 4:5.
它會產出什麼: 一張底片攝影的肖像,帶著準確的復古色彩——真實的顆粒、光暈,以及過期底片的偏色。
為什麼成立: "expired Kodak Portra 400" 是圖片潛空間裡最強的單句風格參照之一——它自帶一整套完整的色調預期。指定偏色的方向("greens shifted slightly toward yellow-olive")可以避開通用的復古顆粒,把偏色引導到過期底片特有的那一種。緊裁加上具體的畫面比例(4:5),產出一張讀起來像真實照片沖印的肖像。
8. 沉浸式環境——降雨(影片)
Ground-level POV inside a glass bus shelter, heavy urban rain, Tokyo residential street, late June 22:00. Camera
holds completely still. Rain streaks down the glass panels in foreground, streetlights smear into vertical bokeh
streaks behind the wet glass. A cyclist passes in the distance — silhouette only, visible for about 2 seconds in
mid-clip. No camera movement. Audio: heavy rain on glass, distant car tyre hiss, one distant motorbike engine
fading right-to-left. No music. 10 seconds.
它會產出什麼: 一段沉浸式的單一視角環境片——當開場鏡頭很強,單獨當氛圍作品也成立。
為什麼成立: "late June 22:00" 一次指定了季節、體感溫度(濕熱的夏雨)和暗度。騎腳踏車的人被安排成一個具體時刻的具體事件("about 2 seconds in mid-clip"),這給了模型一個敘事錨點,又不用它去處理複雜的角色動作。音訊分成三層來寫(玻璃上的雨聲、遠處輪胎的嘶聲、機車聲),通常會產出比單一句「城市雨聲環境音」更講究的音效設計。
在 OmniArt 上跑這些提示詞
八組提示詞都在 OmniArt 創作工作台裡的 Grok Imagine 1.5 上跑得動——不需要另外訂閱 xAI。圖片提示詞(1、5、7)丟進圖片工作台;影片提示詞(2、3、4、6、8)丟進影片工作台的 Grok Imagine。
在 OmniArt 上跑的幾個實務筆記:
- 迭代階段從 480p 開始。 480p 的影片是每秒 10 點數。等結構對了,再拉到 720p(每秒 15 點數)跑最終版。
- 要加長就用延長模式。 氛圍空景(第 3 組)和空拍後拉(第 6 組)都可以用 Grok Imagine 的延長模式再接最多 15 秒——同一個模型,只算接上去那一段的錢。
- 要精準修正就用修改模式。 如果一個結果的光線已經快對了,只有一個元素不對,修改模式讓你用文字描述那個改動,不用重新生成整支片。丟進修改模式之前,來源片請保持 480p——這個模式的輸入上限是 854×480。
- 跨鏡頭的角色一致性: 如果你要生成同一個角色的多顆鏡頭(第 2 組那種),就用參考模式,把大頭照放進
@Image1,並在每一組新提示詞裡重述角色描述。Grok Imagine 1.5 的參考模式,是不靠微調模型就能達到一致性最直接的一條路。
想完整了解 Grok Imagine 全部六種生成模式、成本情境,以及什麼時候該換模型,請看 Grok Imagine 完整指南。想掌握那套能套用到任何影片提示詞上的攝影語彙,電影感 AI 影片提示詞指南值得和這篇一起收藏。
準備好開始創作了嗎?
用 AI 生成精彩內容