This is the blog section. It has three categories: News, Practice Notes, and Releases.
Files in these directories will be listed in reverse chronological order.
This is the multi-page printable view of this section. Click here to print.
This is the blog section. It has three categories: News, Practice Notes, and Releases.
Files in these directories will be listed in reverse chronological order.
IETF 126 met in Vienna, 18–24 July. The collection accessioned its presentation materials as the meeting proceeded; session recordings and minutes follow after adjournment.
In June we promised more on how to browse the newest accessions. The Cypherpunks main list (1992–2016), PacNOG, AfNOG, and W3C www-talk are now indexed and searchable in the Reading Room alongside the rest of the collection. The Africa Internet Summit record (2017–2026) is live as well, with per-year meeting pages and nearly two hundred talk transcripts drawn from the community’s own video archive.
We’ve accessioned the Internet Talk Radio collection — Carl Malamud’s pioneering audio program (1993–1996), 567 items mirrored from archive.org. To make the audio searchable and citable we’re generating transcripts, using the collection itself to correct the errors that speech recognition makes on the period’s names and acronyms. How that works is the subject of our first practice note.
The catalog has added some ninety Works over the past two months. A few worth browsing: the Soviet Coup Internet Archive — the Relcom/Demos traffic that carried news in and out of Moscow during the August 1991 coup, while the broadcast media were seized; ConneXions: The Interoperability Report (1987–1996) and its editorial descendant, the Internet Protocol Journal (1998–present), both edited by Ole Jacobsen; the Internet Experiment Notes (1977–1982), Jon Postel’s RFC-parallel series from the DARPA Internet project; and the Three-Network TCP Demonstration of November 22, 1977, when TCP traffic from a packet-radio van on a Bay Area freeway crossed PRNET, ARPANET, and SATNET on a single end-to-end path.
We’ve soft-launched the IHI Reading Room, a prototype public home for the collection we’ve been assembling: the mailing lists, conference talks, and documents in which the Internet’s operator and standards communities argued the network into existence.
The holdings currently include nearly four million messages of mailing-list correspondence (NANOG, the TCP-IP list, Internet-History, Interesting-People, the regional NOGs, the IETF lists, and many more), tens of thousands of conference presentations and transcripts, and a hand-curated catalog of Works, Authorities, and open research Inquiries, each with a durable ARK identifier under NAAN 26338.
You can search the shelves directly, or ask the librarian a real question and watch it read the primary sources. There’s also an open API and an orientation for AI agents — research assistants are welcome, and their corrections enter the curation queue with honest provenance.
While it’s in prototype, access is by recognition rather than accounts: no passwords, just an arrival link. Get in touch and we’ll send you one. It’s a construction zone, updated daily — have a look, and tell us what’s wrong. That’s what the correction queue is for.
Recent accessions to the IHI collection: the Cypherpunks main list (1992–2016), PacNOG, AfNOG, and W3C www-talk, with Africa Internet Summit materials (2017–2026) staged for processing. More soon on how to browse all of this.
The writeup promised after December’s Barcelona workshop is now available. Envisioning an Internet Data Trust (January 2026) sets out the charter and operational principles for a public-interest institution dedicated to century-scale preservation of the Internet’s historical datasets, and defines the priority tiers that now organize the IHI catalog. Published under CC BY 4.0; comments welcome.
An update on the PingER rescue: a first tranche of historical PingER measurement data has been transferred from SLAC into IHI custody, verified, and accessioned into the IHI catalog (ark:/26338/w/pinger). The project’s servers have since been shut down, and there is no longer anyone at SLAC to ask about the rest — so what we recovered may be all that survives. Decades of end-to-end measurements to and from the Global South, no longer at risk of disappearing entirely.
The Register covered the initiative this week: Internet history is vanishing. Researchers want to save it, by Simon Sharwood — the APRICOT keynote, the replica-hosting model, and the PingER recovery. Nice to see the story reaching a wider audience.
Jim delivered a keynote at APRICOT 2026 in Jakarta: Internet History Initiative: Preserving our Collective Data Legacy (slides). The Internet’s operational exhaust — routing tables, measurements, mailing lists, meeting archives — is disappearing faster than we are preserving it, and the fix is institutional as much as technical. The concrete ask to the room: IXPs, NRENs, and universities willing to host replicas. Get in touch if that could be you.
As part of the Internet Society’s inaugural Pulse Research Week in Barcelona (December 8–11), ISOC Pulse and IHI co-organized a hybrid workshop on December 10: Internet Data Trust: Preserving the Internet’s Historical Data Legacy. More than sixty members of the global Internet measurement community, joined by professional librarians and archivists, shared their institutional experiences and gave feedback on an initial set of community principles. The afternoon covered institutional perspectives on preservation, how to support the custodians we have, and the research horizons that open up if the record survives — working toward a public-interest institution dedicated to century-scale preservation of the Internet’s historical datasets. Thanks to the Internet Society Pulse team for hosting the conversation. A writeup of the ideas is in progress; more soon.
At CAPIF 4 in Almaty, Kazakhstan, Jim presented DNS Without Borders: Uncovering Regional Hubs and Dependencies in K-Root Traffic (slides), work with the RIPE NCC continuing the DNS watersheds theme: where Central Asia’s root-server queries actually flow, and what that says about regional interdependence. A recap of the meeting is on RIPE Labs.
Two talks this spring. At RIPE SEE 13 in Sofia: Assessing the Internet in the SEE Region (slides). And at ZANOG25 in Durban: Contemplating Internet History: Where Do We Go Next? (slides) — the preservation story, told for the operator communities whose mailing lists and meeting archives are the history in question.
Together with Emile Aben of the RIPE NCC, Jim co-authored a RIPE Labs analysis of November’s Baltic Sea submarine cable cuts: A Deep Dive Into the Baltic Sea Cable Cuts. A reminder that today’s measurements are tomorrow’s history: reconstructing events like these depends on the survival of the measurement archives.
Slides from two fall presentations are now available. At the RIPE 89 MAT Working Group in Prague: Measuring and Visualizing DNS Watersheds, on how geography and network proximity shape where DNS queries flow. And at the IRTF GAIA session at IETF 121 in Dublin: Internet History Initiative: Overview and Exhortation, a tour of endangered and extinct measurement projects, and a plea to preserve what remains.
Jim Cowie will be attending DNS-OARC43 and RIPE89 in Prague the week of October 25th, and IETF121 in Dublin the first week in November. Have a story to tell about preserving (or failing to preserve) the Internet’s historical datasets? Stop by and say hi!
After the shutdown of the PingER project at SLAC,
decades of measurement data dating back to the late 1990s were feared lost. We’ve contacted SLAC IT and are in the process of
recovering and archiving this collection for posterity. PingER was notable for its emphasis on creating measurements both
to and from institutions in the Global South, and for its early exploration of possible relationships between national
development and quality of Internet connectivity. More details to follow!
New link to the IXP History Collection, an ongoing project which seeks to record and document global histories of computer networking and internet exchange points (IXPs).
Jim Cowie will be joining Harvard’s Berkman Klein Center as a 2024-2025 Fellow, working with the Harvard Law Library Innovation Lab (LIL) to pursue the Internet History Initiative research themes. Read more here.
The Internet History Initiative has officially joined DNS-OARC . The history of DNS operations is, in a real sense, the history of the Internet. We’re looking forward to connecting with the DNS-OARC community to explore and celebrate this legacy in coming years!
We’ve completed our first offsite preservation mirror of the RIPE RIS historical BGP datasets, all collectors, from inception, and continue to mirror with a 1 day lag. More sources available under Sources.
First sets of links to public data repositories are now
available under Sources. Please
suggest additional public data resources that you’d like to see linked here.
The ARK Alliance has assigned the Internet History Initiative a unique Name Assigning
Authority Number (NAAN) for use in its issuance of unique ARKs,
joining more than 1200 participating institutions
in the ARK Alliance.
Feel free to get in touch if you or your institution would like to learn more about participating in the preservation and curation of the Internet infrastructure’s historical datasets.
Short writeups of methods we use in the collection work, recorded so that others can reuse or improve them.
This note is also available as a typeset PDF.
We recently accessioned the Internet Talk Radio collection (Carl Malamud, 1993–1996; 567 audio items mirrored from archive.org). To make the audio searchable and citable we need transcripts. We ran a single-episode pilot to test a transcription workflow before processing the full collection, using the first “Geek of the Week” episode: an interview with Marshall T. Rose recorded March 31, 1993 (48 minutes).
Automated speech recognition handles ordinary speech in this material well, but systematically mis-renders the specialized vocabulary of the period: acronyms, protocol names, product names, and personal names. The model substitutes phonetically plausible guesses. From the raw output of this episode:
Name errors are the most serious for our purposes, since names are access points. These errors are not correctable by a general-purpose model working alone: nothing in the audio distinguishes “McLaughery” from “McCloghrie” without outside knowledge of who worked with the guest.
The workflow has two stages. Both use information already held in the collection.
Stage 1: speech recognition with a vocabulary prompt. We transcribed the audio with Whisper (large-v3 model, one GPU). Whisper accepts a short free-text prompt that biases its vocabulary. We filled it with the episode metadata (show, guest, interviewer, date) and roughly forty period terms. This reduces, but does not eliminate, the errors above. The output — 741 timestamped segments — is saved unchanged as the raw transcript.
Stage 2: constrained correction using a brief derived from the collection. We then passed the transcript in chunks to a language model (Claude Sonnet) with a correction brief and narrow instructions: correct only mis-transcribed technical terms, acronyms, and proper nouns that the brief supports; do not rewrite, paraphrase, or improve the speech; do not alter timestamps.
The brief (about 1,100 characters) was assembled from two collection sources:
The corrected output is saved as a second file alongside the raw one. Both are kept permanently: the raw transcript is the preservation copy; the corrected transcript is the access copy; a line-by-line comparison of the two shows every change made.
The correction stage changed 41 of 741 segments. Examples, with timestamps into the recording:
| Time | Raw output | Corrected |
|---|---|---|
| 0:56 | the author of ISO-DE | the author of ISODE |
| 6:02 | Marty Shostal of PSI | Marty Schoffstall of PSI |
| 7:33 | X400 (and 5 further occurrences) | X.400 |
| 17:15 | a call from the IEPF (and 5 further) | a call from the IETF |
| 17:42 | Keith McLaughery | Keith McCloghrie |
We reviewed the full diff. All 41 changes were within the instructed scope (terms, acronyms, proper nouns); none rewrote the wording of the speech.
Processing cost for the episode: about 9 minutes of GPU time for transcription and 8 language-model calls for correction. Extrapolated to the full 567-episode collection: roughly 10 GPU-hours and a few thousand model calls, which is within our normal batch budgets.
The workflow assumes nothing specific to internet history. It requires: (a) an ASR system that accepts a vocabulary prompt; (b) a language model that will follow narrow correction instructions; and (c) collection materials contemporaneous with the recordings — correspondence, minutes, newsletters, an authority file — from which to build the brief. For archives that hold such materials, the marginal effort per recording is small: the brief for this episode was assembled automatically from existing records.
Contact: the IHI Reading Room. Tooling: Whisper large-v3 with initial_prompt; Claude Sonnet for correction; brief assembly scripted against our catalog API and message index.