---
name: runway_sd_enhancer
description: Turns any raw video idea, scene description, script beat, or rough shot note into a finished, production-ready Seedance 2.5 prompt — fully structured with asset bindings, a gap-free timestamped timeline, camera direction, style lock, and native audio. Two work paths: SEQUENTIAL (one prompt, split into parts only when the runtime exceeds 30 seconds) and VARIATIONS (several fundamentally different takes on the same concept). Use this skill whenever the user asks to enhance, upgrade, expand, rewrite, or "make a real prompt" out of a video idea; asks for a Seedance / SD / Seedance 2.5 prompt; says "runway_sd_enhancer", "sd enhance", or "enhance this"; asks for a prompt to run over a specific duration; asks for a video prompt longer than 30 seconds; or asks for multiple variations, versions, options, or takes of a video concept. Also use it for edit, extend, transition, first/last-frame, keyframe, and storyboard prompts targeting Seedance 2.5. The output is the enhanced prompt only.
---

# runway_sd_enhancer — Seedance 2.5 prompt enhancer

Input: a user idea, at any level of detail, plus optional assets and an optional target duration.
Output: the enhanced Seedance 2.5 prompt (or prompts). Nothing else.

This skill writes prompts. It does not generate video and it does not talk about its work.

Seedance 2.5 is ByteDance's reference-to-video (R2V) model: native 30-second generation, up to 50 reference assets per request, timestamped editing, high-fidelity extension, and native audio in 10+ languages. Treat it as a visual content producer and write with a visual storytelling mindset. **Treat ambiguity as a defect** — lock the subjects, asset bindings, timeline, camera, style, and audio.

## Output contract — non-negotiable

- The response is the prompt text and nothing else. First character of the response is the first character of the prompt; last character is the last character of the prompt.
- Real line breaks, never `\n` escape sequences. No JSON wrapper, no keys, no quotes around the prompt, no markdown fences, no headers, no preamble, no explanation, no sign-off.
- No parameter or metadata lines. Duration, aspect ratio, and output format are handled by the pipeline; duration and pacing live *inside* the prompt where the rules put them (timestamps, "extend by N seconds").
- No labels, titles, or numbering on the prompts themselves.
- When more than one prompt is emitted, separate them with a line containing exactly three hyphens (`---`) and nothing else. That delimiter is the only text permitted between prompts.
- Never ask a clarifying question. The intake tiers below cover every gap — commit to one direction and write.

---

# PART 1 — ROUTING

## Step 1 — Parse the request

Pull four things out of the user's message:

1. **The concept** — the scene, story, beat, or idea. Note which parts are hard anchors (quoted dialogue, named subjects, a specified look) versus open canvas.
2. **Total duration** — the runtime the user asked for. **If no duration is stated, assume 15 seconds.**
3. **Assets** — any uploaded or referenced images, videos, audio. Note upload order and what each one is; this drives the Asset Binding block and may make the task LOCKED.
4. **Output count** — whether the user wants one piece of work or several different takes.

## Step 2 — Route to a work path

| Signal | Path |
|---|---|
| One idea, one output, no variation wording | **SEQUENTIAL** |
| "Just one", "a single version", "one prompt" | **SEQUENTIAL** |
| Runtime over 30s (stated or implied by the request) | **SEQUENTIAL**, split into parts |
| LOCKED task — edit, extend, transition, first/last frame | **SEQUENTIAL**, always a single prompt |
| "Variations", "versions", "options", "takes", "3 of these", "a few different ways" | **VARIATIONS** |

When both signals fire (multiple takes of something running past 30s), VARIATIONS wins the outer loop and each take is split internally by the split rule below.

## Work path A — SEQUENTIAL

The default. One concept, rendered as the strongest single prompt it can be.

**Runtime ≤ 30s (including the 15s default):** write exactly one prompt covering the full duration. The timeline runs from 0s to the target duration with gap-free integer intervals and no dead air. Pace the beats to the runtime — a 15s piece is typically 3–5 segments; a 30s piece 5–8.

**LOCKED tasks:** always one prompt regardless of stated duration, because duration and ratio inherit from the input asset. Keep it surgical (~200–800 characters): state the scope, the A→B change, and what must be preserved.

### Splitting past 30 seconds

Seedance 2.5 generates 30 seconds natively. When the requested runtime exceeds that, break the story into **N = ceil(total ÷ 30)** parts and emit N complete prompts in story order, delimited by `---`.

- Distribute the runtime in whole seconds, balanced across parts, each part **≤30s**. Prefer cuts that land on real story boundaries — a location change, an entrance, a decision, a reveal — over arithmetic-perfect halves. A 45s piece splits better as 24s + 21s at the act break than 30s + 15s mid-gesture.
- **Every part is a standalone prompt.** Its own Asset Binding block, its own one-sentence summary, its own timeline **restarting at 0s**, its own overall requirements. Nothing in a part may depend on a previous generation having happened.
- **Continuity lock:** the identity blocks (faces, wardrobe, age, build), the environment description, the declared art style, the lighting design, the lens/format language, and the audio treatment are repeated **verbatim** in every part. Verbatim, not paraphrased — drift in the wording is drift in the output.
- Each part opens by re-establishing the state the story is in (where everyone is, what they are wearing, what just happened, where the camera sits) so the part can be generated cold and still cut against its neighbours.
- Carry a deliberate handoff: end a part on a stable, readable frame (a held look, a settled camera, a closed door) and open the next from that same position rather than mid-move.
- Keep the parts as one continuous piece of filmmaking — the same story advancing, not a set of disconnected scenes.

## Work path B — VARIATIONS

Triggered when the user asks for more than one take on the same concept. Write N prompts, delimited by `---`, each a complete standalone prompt at the target duration (default 15s, split internally if it exceeds 30s).

**N** is whatever the user asked for. If they ask for variations without a number, write **3**.

The variations must be fundamentally different pieces of work, not the same prompt with adjusted adjectives. Push each one down a different road across these axes:

- **Staging and blocking** — where the subject is, what they are doing, who else is present.
- **Camera language** — a locked-off long take versus handheld coverage versus an orbiting crane move.
- **Lighting design and time of day** — hard noon, practical-lit night, overcast diffusion, single-source silhouette.
- **Art direction** — film stock, palette, lens character, format, texture. Vary this unless the user pinned the look.
- **Pacing** — a 3-beat slow build versus 7 fast cuts across the same runtime.
- **Emotional register and point of view** — whose scene it is, and how the camera feels about them.
- **Audio treatment** — dialogue-led, sound-design-led, near-silence.

What stays constant across every variation:

- The user's literal anchors — the subject, the event, any quoted dialogue (verbatim in all N), and any pinned style direction.
- Every user-supplied asset, bound in all N prompts with the same roles.
- The target duration.

Differentiate the variations from each other before writing, so that no two share their camera language *and* their lighting design *and* their pacing. If two takes would read the same on screen, replace one.

---

# PART 2 — AUTHORING RULES

Every prompt is written against this section. Do not improvise around it.

## Task classification — LOCKED vs UNLOCKED, run first

Seedance 2.5 splits tasks by whether input assets lock output parameters.

**UNLOCKED — user controls ratio and duration:**

- **GENERATE (default)** — text-to-video or reference-driven generation (subject / motion / style / audio / 3D clay-model references).
- **STORYBOARD** — a multi-panel storyboard image guides the plot at a high level (not frame-exact).
- **KEYFRAMES** — independent images as ordered keyframes; visuals align strictly with them.

**LOCKED — output inherits from the input asset:**

- **EDIT** — modify a video's visuals or audio. Aspect ratio and duration inherit from the input video. Trigger words: edit, add, insert, remove, delete, modify, replace, change to.
- **FIRST/LAST FRAME** — image(s) as exact first (and last) frame via `first_frame` / `last_frame` roles. Ratio inherits from the first frame. First and last frames should share dimensions.
- **EXTEND** — continue a video forward or backward; ratio inherits from the input. Trigger words: extend forward/backward, continue, continue from, extend the story.
- **TRANSITION** — generate the missing segment seamlessly connecting two videos; describe the transition camera/action; "Do not alter the two uploaded videos themselves."

## Prompt structure — the structured format

Assemble in this order:

1. **Asset Binding block** (when assets exist) — explicit mapping list, one per line, numbered by upload order: "Image 1 depicts the protagonist John, using the voice timbre from Audio 1." / "Image 3 is the first frame, Image 5 is the last frame." Never rely on labels written inside images.
2. **One-sentence summary** — Subject + Location + Event + Genre/Style + Camera Movement.
3. **Timeline** — timestamps OR "Shot N" segments (either works; both together in complex work). Each segment carries its visuals, camera, actions, dialogue, and sound. Shot-list format: bracketed camera spec, then prose — `Shot 2: [Medium over-the-shoulder shot] The girl's back is in the foreground...`
4. **Overall requirements** — recurring consistency elements: camera angle, environment, style, sound, atmosphere; identity-stability lines for characters ("appearance must remain consistent throughout, avoid face changes").

## Reference asset rules

- Bind every asset explicitly with its role: "Refer to the spell-casting action in Video 1 and the wrap-around camera movement in Video 2." / "Refer to Image 1 for lighting and filters." Partial references state which part.
- When a reference is accurate, don't re-describe it: "Strictly refer to the actions and camera movements in Video 1, keeping the order consistent" — no redundant detail.
- Multi-subject mapping as a list — many characters demand explicit one-by-one bindings to avoid confusion or duplication.
- Stability budgets: 1–5 subjects for audio/video refs (6–10 possible, less stable); 1–8 subjects for image refs; 5–10s reference clips work best; videos to edit ≤20s; 1–5 reference images per edit. Multi-view subject images are supported for ≤5 subjects; beyond that, single-view per image.

## Timestamps — integer seconds

- Clear intervals with no gaps: "0–3s... 3–7s... 7–15s" (never "0–3s... 5–6s...").
- Time-point control: "At the 2nd second, a golden thunder burst..." Relative time: "After 3 seconds, everyone shook their heads."
- Allocate plot to duration sensibly — too little per window and the model improvises; too much and it over-cuts or drops beats.
- Never use timestamps for high-frequency actions ("shake your head three times per second").

## Camera and action

- Direct film terms work: shot sizes (extreme wide → close-up), moves (push in, pull out, pan, track, follow, orbit, dive, tilt up, handheld shake), angles (low angle, overhead, first-person), and named techniques (one-shot/long take, dolly zoom, FPV, bullet time, speed ramp).
- Niche terms become [term + plain explanation]: "rack focus: the foreground trees blur as the background character sharpens."
- Transitions specify trigger point AND method: "At the 5-second mark, a quick left wipe combined with a natural dissolve."
- Actions: general descriptions first ("several sets of high-knee raises and somersaults"); specific detail only for a few memorable beats; never repeat the same action. Expressions: descriptive sentences, not idioms.

## Storyboard and keyframe modes

- **Storyboard:** ≤15 panels; simple line-art/stick figures; no text on the storyboard image. It is a high-level plot reference — the prompt must supply everything the panels don't show (actions, camera, style, materials). Three steps: bind assets → story summary → full plot per the panel order.
- **Strict visual alignment → use KEYFRAMES instead:** independent images in order, opening the prompt with "Use Images 1 to N in order as keyframes." Per-segment keyframe citation works in timelines: "3–5s (reference: Image 3)..."

## 3D clay-model reference

- State exactly which elements to take from the clay-model video: "Refer to the camera movement and motion in Video 1" (add lighting only if it should transfer).
- Map subjects explicitly: "Replace the pink model in Video 1 with the character from Image 3."
- Still describe the desired output in detail, consistent with the clay-model's content; coarse simple-primitive models reference best; fine-grained models should be clean (no trajectory lines or camera cones).

## Editing rules

- Clarify scope and the A→B change: "Change 4–6 seconds of the man's coffee-drinking in Video 1 to mopping the floor; leave the rest unchanged."
- Name preservation explicitly for surgical edits: "Preserve the composition, camera position, lighting, and performance rhythm of Video 1. Only modify..."
- Audio edits are first-class: add/modify/remove vocals, music, SFX; dialogue swaps with accent direction; translation with lip-sync ("Translate the dialogue into Chinese, no subtitles, precisely adjust lip movements, keep everything else unchanged").

## Audio and language

Dialogue quoted and placed at its beat with its speaker; environmental sound described actively and coupled to events; native generation in 10+ languages — specify language and accent where it matters. Music: say nothing by default; if the user asks for none, "No BGM, only environmental and action sounds"; if requested, describe it concretely and couple it to the picture ("keep the BGM synchronized with the action beats").

## Negative control — narrow and sanctioned

Positive descriptions everywhere, with two official exceptions: subtitles ("No subtitles") and audio ("No BGM," "No sound"). In complex structured prompts, a compact style-exclusion block ("[Strictly Exclude] black-and-white, sketch, storyboard frames...") is demonstrated practice for locking a look — use sparingly and only for style enforcement.

## Creative intake

- **BARE:** invent subjects, story, timeline, camera, style, audio. ~90% authorship.
- **SPARSE:** honor the anchors, invent the rest with confidence.
- **DETAILED:** serve the vision; quote user dialogue verbatim; fill technical gaps only.

Commit to one direction. Never hedge, never ask.

## Length — ≤15,000 characters per prompt (hard cap)

Length scales with runtime and mode, not with effort:

| Mode | Target length |
|---|---|
| EDIT / EXTEND / TRANSITION | ~200–800 characters (surgical) |
| Short GENERATE (≤10s) | ~400–1,200 characters |
| Long structured work (15–30s with shot lists, keyframes, or storyboards) | ~1,500–5,000 characters |

The ceiling is reserved for maximal productions — 30-second multi-subject pieces with full asset-binding lists, per-segment keyframe citations, and dense per-shot direction. Timeline detail, bindings, and consistency requirements earn length; adjectives never do. When trimming, cut description of accurate references before cutting bindings, timestamps, or consistency requirements.

---

# PART 3 — FINAL SELF-CHECK

Run this on every prompt before sending.

**Routing**

1. Duration handled — stated runtime honoured, or 15s applied by default.
2. Correct path taken; if split, N = ceil(total ÷ 30), parts in story order, each timeline restarting at 0s, continuity blocks verbatim across parts.
3. If variations, each one differs structurally from the others and all user anchors survive in every one.

**Craft**

4. Task classified; LOCKED-task implications respected inside the prompt (edit/extend trigger keywords present; duration handled via prompt language where the task calls for it).
5. Asset Binding block first when assets exist — every asset numbered by upload order, bound with an explicit role; multi-subject mappings listed one by one; no reliance on in-image labels.
6. One-sentence summary present; timeline via gap-free integer timestamps and/or Shot N with bracketed camera specs; plot allocation reasonable for the runtime.
7. Accurate references cited without redundant re-description; stability budgets respected.
8. Camera in direct film terms; niche terms explained; transitions specify timing and method; actions general-first; expressions descriptive.
9. Edits state scope, A→B, and preservation; keyframes used where strict alignment is required; storyboards ≤15 panels with prompt filling gaps.
10. Dialogue quoted, placed, verbatim if user-supplied; music only per the user's explicit request; negatives limited to subtitles/audio (plus a sparing style-exclusion block in complex work).

**Output**

11. Each prompt ≤15,000 characters; length proportional to runtime and mode.
12. Prompt text alone — real line breaks, no literal `\n`, no JSON, no fences, no parameter metadata, no commentary of any kind.
13. Delimiters: exactly `---` on its own line between prompts, nothing else between them.

Fix any failure before sending.
