Most viewers decide in under three seconds. A working hook does three jobs fast: interrupt the scroll (0–1s), promise something specific (1–2s), and start delivering it (2–3s). The reliable hook types — open loops, bold claims, direct address, pattern breaks — all fit that anatomy.
Creators polish minute two of videos that lose the audience in second one. The first three seconds aren't part of the video — for most viewers, they ARE the video.
Second one has one job: stop the thumb. Movement, an unexpected frame, a claim that can't be ignored, a question mid-thought. Second two makes the promise: what, specifically, does staying get me? Second three starts paying it off — not 'coming up later,' but the first piece of the answer, now.
Videos that do all three keep their test batch; videos that spend those seconds on intros, logos, or 'hey guys, welcome back' hand the algorithm its exit data before anything happened.
The reliable families, each transferable to any niche: the open loop ('nobody talks about the third one'), the bold claim ('this replaced my entire routine'), direct address ('if your skin does this, watch'), the pattern break (starting mid-action, mid-sentence, mid-result), the enemy hook ('everything you've heard about X is backwards'), and the result-first hook (show the after, then rewind).
None of these are magic words — they're structures. The specificity is what carries them: 'this $9 product' beats 'this product,' 'in four days' beats 'quickly.' Vague hooks read as ads; specific ones read as information.
The fastest hook education is studying what already won in YOUR niche. Find the outliers — the videos that clearly beat their creator's average — and transcribe their openings word for word. (Pulling a video's speech to text takes a minute with a transcriber.) Then look past the words to the move: what job did second one do? Where was the promise? When did delivery start?
Build a swipe file of ten transcribed openings from your niche and the patterns become impossible to miss — and legitimately yours to reuse, because structure isn't plagiarism.
On slideshows and captioned videos, the FIRST FRAME is the hook: the text on slide one does the interrupting and promising before any audio matters. The same anatomy applies — specific claim, visible instantly, positioned where the UI can't cover it (run it through a safe-zone check; a hook under the caption bar is a hook nobody read).
And on any video: assume sound-off. If the hook only exists in audio, it doesn't exist for the muted majority — put it on screen too.
The work happens in the first three seconds: interrupt, promise, and begin delivering. Everything after is retention, not hook.
Whichever fits the content honestly — open loops, bold claims, direct address, and result-first all work when the promise is specific and kept.
Yes — slide one's text IS the hook. Make it a specific claim and keep it out of the UI-covered zones.
Copy the structure, never the sentences. Formats are shared property; scripts aren't.