Condé Nast speeds up video archive search with multimodal AI search on Amazon Bedrock
Condé Nast introduced multimodal search across its archive of 140 000+ videos using Amazon Bedrock, Amazon OpenSearch Service, and the TwelveLabs Marengo model. According to the company, this reduced footage search time from 250 minutes to less than 2 minutes per task.
Publishing house Condé Nast (Vogue, GQ, Vanity Fair, Wired) deployed, in collaboration with the AWS Generative AI Innovation Center, a multimodal search solution for its archive of more than 140 000 videos. According to the company, editors previously searched for needed footage for an average of 250 minutes per task, because existing tools could only search file names and captions, not the actual video content. The new solution, according to Condé Nast, reduced this time to less than 2 minutes per task.
The system is built on Amazon Bedrock and Amazon OpenSearch Service and uses the TwelveLabs Marengo embedding model, which, according to the article, can jointly encode the visual, audio, and textual (transcript) signals of video into a single vector space. This allows searching based on the intent of a natural-language query instead of exact keywords in metadata.
The architecture separates two layers: asynchronous ingestion and indexing of embeddings (computationally intensive video processing) and a synchronous layer for serving search queries. According to Condé Nast, this separation allows independent scaling of both parts and was especially crucial during the initial backfill processing of the entire 140 000-video archive, when search remained available even during ongoing reindexing.
The source article is a partial overview of the solution and contains only an excerpt from a more extensive technical breakdown by AWS; for details on the architecture and other measured results, see the source article.
Why it matters
This case shows a concrete, measurable benefit of multimodal embeddings for companies with large archives of unstructured content (video, audio) — instead of searching metadata, one can search by the actual content of the footage, audio, and spoken word. For media companies, marketing departments, or archives, this is a relevant pattern for reducing operational delay in finding material and limiting reliance on individual employees' institutional memory.
Relevant practical impact
What this means
For a business
A company with a large archive of video or other multimedia content can replace searching by file names and metadata with semantic search across image, audio, and transcript, which according to Condé Nast significantly reduces the time needed to find a specific shot and reduces reliance on a specific employee knowing where the material is located.
ProductivityCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.