Reconstructing the missing decks

A practice note on recovering slide evidence from video-only conference talks

This note is also available as a typeset PDF.

Of the 11,688 video recordings in the collection’s talk holdings, 11,240 — ninety-six percent — have no accompanying slide deck. The deck is usually where the substance is: the numbers, the network diagrams, the timeline the speaker is narrating. A video without its deck is searchable only by its title. The archive knows a talk happened, but not what it said.

The method has two steps. The first is a hunt: a per-talk web search for the original deck, escalating to event-page fetches and Wayback lookups. When the talk comes off the conference circuit this often works on the first query, and it doubles as a metadata audit — one recovered deck exposed a misspelled presenter name in our own catalog. When the source is a streaming channel the hunt mostly reports honest absence: of forty talks from one such channel, one deck found, twelve confident verdicts that no deck ever existed (panels, welcome remarks, an award induction), and twenty-seven misses. A well-evidenced “probably no deck” is a finished piece of work, not a failure.

Failing the hunt, we reconstruct the deck from the video. Sample a frame every five seconds, cluster the frames into slide-states by perceptual hash, keep the sharpest frame of each state, and drop states too short to be slides — those are camera cuts. Each surviving page records the timestamp span over which it was visible, so every reconstructed page links back to the second of video where the speaker is saying it. About four minutes of CPU per talk.

A vision pass over 1,023 reconstructed pages found the population divides three ways: 282 switched-feed slides, where the production cut to the deck; 347 projections in the scene, where the room camera happens to see the screen; and 377 pages with no slide at all. The projection class was the surprise — talks whose metadata said “panel” turned out to carry legible keystoned decks on the room’s screen.

The finding we did not expect: for the one talk where we hold both the found deck and the reconstruction, the reconstruction is the better artifact. The archived PDF is a print-resolution two-up handout; the video’s switched feed shows each full slide at 720p. Reconstruction is not only a fallback for missing decks. For switched-feed recordings it can be the archive’s best edition, with the found deck serving as the authority on slide count and order.

A reconstructed deck is a derived artifact and must never pass as the original. Each carries the source video’s URL and hash, the full method string, per-page timestamp spans, the model that transcribed each page, and a verdict distinguishing a reconstruction from a confirmed absence of slides. In the catalog they bind at receipt tier, honest about what they are.