Subchapter 14.1
references/aesthetic-principles.mdMarkdown8 KBView on GitHub
The 18 rules that separate “designed” caption work from “preset-generated” caption work. This is the single most important reference in the skill — every plan.json and every Standard-mode HTML should be checked against these before committing.
The whole competitive thesis: every AI caption tool in 2026 (Veed, Submagic, Opus Clip, Captions.ai, CapCut) is a . They hand the user a box of crayons. This skill is a — it exercises judgment. Beat them on taste, not feature count.
Also bundled
GitignoreCaptions exist to serve the face, not compete with it. If the caption draws the eye before the speaker, the caption has failed. Before shipping any plan.json, ask: “does the caption or the subject read first?” If caption wins, shrink, dim, or move it.
World-class caption design puts text into the scene. Letters that pass behind a shoulder, mic, or head feel diegetic. Letters that float uniformly above the lower-third feel like PowerPoint. Use the matte pipeline — that’s the moat over every competitor.
A white box behind text is a failure of taste. Priority order:
mix-blend-mode: overlay or screen picking up scene luminanceFixed-bottom captions are monotone. Caption zone shifts with shot:
Hierarchy lives in weight (e.g., 500 → 800), not in font. Mixing Montserrat + script + serif in one clip is the #1 amateur tell. Ship Inter, SF Pro, Söhne, GT America, Aktiv Grotesk, or Neue Haas Grotesk with ≥5 weights and compose with weight + size.
Display-size (>40pt) wants negative tracking (-10 to -30 units, or -0.015em to -0.035em). Body-size (14–20pt equivalent) wants positive tracking (+5 to +15, or +0.005em to +0.015em). Apple SF Pro’s optical-size model is the reference. Submagic defaults do the opposite — they look cheap because of it.
Italics are a print convention for flow inside a paragraph. On 24fps motion they read as “tilted” not “stressed”. Emphasize with weight (extrabold), color (single accent), or size (1.3–1.6×). Italics allowed only for:
Pick one saturated accent per video for keyword highlights. Hormozi’s yellow+green+red works for him because his content is already loud. Cinematic = single accent + white/bone/charcoal. Default palette:
#F5EFE6 on dark#1A1A1A on lightletter-spacing, filter:blur, font-weight animations cause inline-block reflow → line-jumps. Animate translateY, scale, opacity, clip-path. This is locked from the embedded-captions debugging.
BBC reading speed is 160–180 wpm (0.33–0.38s/word). Cinematic feel wants more breathing room — 200–220ms stagger per word, hold full phrase 0.4–0.6s before exit. Entries under 150ms read as frantic.
Same phrase, different stagger, totally different feel:
| Stagger | Feel |
|---|---|
| 40ms | machine-gun, urgent, TikTok-hook |
| 80ms | conversational, default |
| 150ms | deliberate, documentary |
| 250ms+ | poetic, ceremonial |
Pick from content tone, not default to one value.
Every word bolded = no emphasis. Structure:
Every-word-same-way = eye adapts, stops registering motion. Plant a rhythm-break every ~30s:
This is what separates Submagic-preset work from something designed.
Transcribe everything. Display 70–85%. Remove:
Editorial judgment — no existing AI caption tool does this. It’s pure upside.
Chunk at natural pauses ≥ 250ms. A caption spanning a breath-break feels wrong. Whisper word-level timestamps make this trivial.
/hyperframes-studio (Safe zones) for wide and vertical
framings; that skill owns the values.Bake into the layout solver. Never eyeball.
Black bars on 9:16 from 16:9 source? Those bars are the caption home. 2.35:1 cinematic frame on 16:9 export? Serif quotation in the letterbox reads as documentary.
These are the “what to caption” decisions. No existing tool exercises them.
[laughs] / [sighs] — those are accessibility captions. For aesthetic captions, they pollute the frame.The agent’s own pre-render pass. Flag violations:
Any “yes” to a violation question → regenerate the affected segment.