Video Search
Search Footage by What's Said vs. What's Shown
Some moments are only findable by what's said out loud; others only by what's on screen. A footage search system needs both layers, not just a transcript.
Two separate signals, two separate gaps
Transcript search finds a moment because of what someone said — useful for interviews, podcasts, and talking-head footage, but blind to anything silent: a reaction shot, an establishing shot, a product close-up. Visual search finds a moment because of what's on screen — useful for exactly the footage transcript search misses, but blind to spoken content that isn't visually distinctive.
A tool that only does one of these will consistently miss moments that only exist in the other layer. "Find the shot where he's nodding" is invisible to transcript search. "Find where she mentions the deadline" is invisible to visual-only search.
Why both layers need to be indexed together
The moments that make an edit work are often a combination: what's being said, over what's being shown. Cutting a voiceover to the right visual, or finding a reaction to match a line of dialogue, requires the system to reason across both layers at once — not run two separate searches and merge the results by hand.
How Broll handles it
Broll indexes both dialogue and visual content for every clip, so a search or an editorial decision can draw on either signal — or both together — without you having to know in advance which one holds the moment you need.