How Condé Nast built multimodal video discovery with Amazon Bedrock
Condé Nast’s editorial teams had no fast way to do multimodal video discovery. They were spending an average of 250 minutes per content discovery task, manually scrubbing through a library of more than 140,000 videos. They relied on titles and descriptions to find relevant clips. In a media environment where speed-to-market directly determines revenue capture, this process created measurable operational drag across brands such as Vogue, GQ, Vanity Fair, and Wired.
The core problem was structural: Existing search tools can’t look inside video content. Teams depended on institutional knowledge to locate assets, creating single points of failure when specific individuals were unavailable. Meanwhile, underutilized content sat in the archive undiscoverable because no keyword in a title or description connected it to the queries editors were actually running.
To solve this, Condé Nast partnered with the AWS Generative AI Innovation Center (GenAIIC) to build an AI-powered multimodal video discovery solution. Built on Amazon Bedrock and Amazon OpenSearch Service, the solution runs intent-based semantic search across video transcripts, visual elements, and audio. The team selected the TwelveLabs Marengo embedding model for its native ability to jointly encode visual, audio, and transcript signals. Marengo powers all five capabilities described in the following section. This reduced discovery time from 250 minutes to under 2 minutes per task.
In this post, we describe the architecture, explain why we chose specific technology choices, and share the business outcomes the solution delivered.
Why semantic search (and why this architecture)
When the team scoped the problem, two constraints shaped the solution design.
First, keyword search was fundamentally insufficient. Editorial teams don’t search for “yoga_tutorial_march_2024.mp4.” Instead, they search for “beginner yoga content with calming backgrounds” or “behind-the-scenes fashion week moments.” The search layer needed to understand intent, not match strings. This pointed directly to vector embeddings that capture semantic meaning across modalities (visual, audio, and transcript). Human-authored metadata alone could not provide that level of understanding.
Second, the video library was large enough (over 140,000 videos) that any solution needed to separate the expensive, compute-heavy work of generating embeddings from the low-latency work of serving search results. A monolithic architecture would force a tradeoff between ingestion throughput and query responsiveness. Decoupling them meant each plane could scale, fail, and evolve independently. This decision proved essential during backfill processing.
These two constraints led to the solution’s core design: multimodal vector embeddings generated by the TwelveLabs Marengo model on Amazon Bedrock, indexed in Amazon OpenSearch Service, and served through a purpose-built query tier. This tier converts natural language into vector searches and returns precise timestamps. We accessed Marengo through Amazon Bedrock because Bedrock offers a range of foundation models (FMs) through a single API. It also applies AWS governance controls, including AWS Identity and Access Management (IAM) for access, Amazon Virtual Private Cloud (Amazon VPC) for network isolation, and AWS CloudTrail for auditability. That combination let the team adopt a specialized embedding model without building or operating its own model-serving infrastructure. For the vector index, Amazon OpenSearch Service provides managed k-nearest neighbor (k-NN) search with multi-AZ replication and metadata filtering for hybrid queries. This let the team run low-latency similarity search across more than 140,000 videos without managing the underlying search cluster.
The solution gives editorial teams a fundamentally different way to interact with their video library. The following capabilities are available through the solution:
The following diagram illustrates the high-level architecture, showing the ingestion and indexing plane and the query and serving plane.
Figure 1: High-level architecture of the multimodal video discovery solution, showing the ingestion and indexing plane and the query and serving plane.
The solution consists of two decoupled planes: an asynchronous ingestion pipeline that makes videos searchable, and a synchronous serving tier that handles user queries.
When a new video is uploaded, the following sequence executes:
The pipeline is event-driven and orchestrated end to end. This gives the team per-step retry logic, parallel processing across chunks, and full traceability for debugging. These were capabilities the team relied on heavily during the initial 140,000-video backfill.
When a user performs a search, the following sequence executes:
The solution is multi-AZ throughout. The following components are deployed for resilience:
Condé Nast conducted a benchmarking workshop in May 2026 to quantify impact. The following results reflect Condé Nast’s measurements from that workshop.
“The AWS team, along with the GenAI and TwelveLabs attendees, helped us explain the business case clearly and concisely. I am optimistic about notching more progress towards our workflow goals in the near future.”
— Billy Keenly, Global Senior Director, Creative Optimization, Condé Nast
Building this solution across an over 140,000 video library surfaced practical lessons that go beyond what architecture diagrams capture.
Early user research shaped the entire embedding and query design. The team interviewed editorial staff to understand how they describe content in their own words. Those conversations defined the right abstraction level for semantic search. Without that grounding, the system risked optimizing for queries no one actually types.
Separating ingestion from serving proved essential at this scale. Each plane could evolve, scale, and fail independently. The solution has been running in production for six months, and production search stayed available while the ingestion pipeline reprocessed backfill content or received updates.
Finding the right video segment length required deliberate experimentation. Short segments lost context. Long segments diluted the semantic signal. Iterative benchmarking against real editorial queries drove the team toward a duration that balanced precision and recall.
Asynchronous embedding generation was non-negotiable at this scale. Synchronous calls to the embedding model would have created bottlenecks across over 140,000 videos. Asynchronous invocation through AWS Step Functions allowed the pipeline to process the backlog without blocking, and it remains the pattern for ongoing ingestion.
Building in multi-AZ availability from day one avoided costly retrofits. For a solution that editorial teams rely on throughout the workday, even brief outages translate directly to lost productivity. It’s cheaper to build this in from the start than to add it later.
This post described how Condé Nast and AWS built a multimodal video intelligence solution using Amazon Bedrock, Amazon OpenSearch Service, and TwelveLabs Marengo. The solution replaces metadata-only search with intent-based semantic search across visual, audio, and transcript modalities. The result: a 99.2 percent reduction in discovery time and an estimated $800,000 in annual operational savings across an over 140,000 video library.
This pattern applies to organizations managing large video libraries, including broadcasters, streaming services, sports leagues, and enterprise media teams. The decoupled architecture adapts to different embedding models and content types as multimodal AI capabilities evolve.
To get started with multimodal video discovery on AWS, visit Amazon Bedrock and Amazon OpenSearch Service. To learn how the AWS Generative AI Innovation Center can help your team build AI-powered solutions, visit the AWS Generative AI Innovation Center page.
We would like to thank the following contributors for their work on this solution: Shinan Zhang, Xiaoye Qian, Jessica Malca, Elvis Joseph, and Jason Janetzke.
Related Stories
AI News
Trump signs executive order to launch AI
13 minutes ago
AI News
Trump vows to never 'stifle' AI as he gathers with tech executives urging caution
13 minutes ago
AI News
The human touch in an AI
40 minutes ago
AI News
Ex
41 minutes ago
AI News
As AI fears ‘spread’, Donald Trump to meet CEOs of Meta, Google, Nvidia, Anthropic and other tech compani
43 minutes ago
AI News
Prompt engineering fundamentals for Amazon Quick
1 hour ago
AI News
60% of leaders mention AI at UN General Assembly speech
1 hour ago
AI News
Exploring the ethics of artificial intelligence
1 hour ago