Projects > Wetpixel Archive
Wetpixel Archive: Personal preservation archive of Wetpixel.com: 412,000 forum posts, 7,800 articles, and 80 hours of podcast, structured and browsable offline.

Wetpixel Archive

Personal preservation archive of Wetpixel.com: 412,000 forum posts, 7,800 articles, and 80 hours of podcast, structured and browsable offline.

Year
Jun 2026
Status
Live
Platforms
Home server
Type
Pipeline, Automation
Role
Sole developer
Links
Wetpixel Wiki ↗Alex Mustard's BSoUP post ↗Wiki post on Waterpixels ↗

I ran Wetpixel for many years; the current live site is at risk of disappearing any day. This is the complete archive: every article, forum thread, comment, and podcast episode, preserved before it was lost, then read end to end by AI to write a wiki charting the first 20 years of digital underwater photography.

A personal preservation archive of Wetpixel.com, built by Eric Cheng, who ran the site for years, to keep the community's history intact. It saves a complete copy of the main site and then turns that copy into everything else offline: a searchable dataset of 7,872 articles and 1,526 news items, a browsable offline replica of the site, the community's comment threads matched back to the articles they belong to, and downloaded videos.

The discussion forum was captured in full: 412,725 posts across 64,016 threads from 14,960 members, spanning December 2001 through May 2026. A separate effort archives 307 Wetpixel Live episodes from YouTube, including audio, video, and written transcripts, totaling over 80 hours of content.

The site was crawled once and everything else built from that single copy, to keep the load on the live server low. Images that had gone missing were recovered from an older backup (about 1.08 GB) and merged back in. The finished archive is 158 GB on disk.

On top of the archive sits the Wetpixel Wiki, published at wetpixel.echeng.com: an AI-written history of the site and of the first 20 years of digital underwater photography, distilled from the full archive into nearly 300 searchable pages covering the year-by-year timeline, notable people, gear, techniques, events, companies, and dive destinations.

Features

  • Every forum conversation preserved: 412,725 posts across 64,016 threads from 14,960 members, spanning December 2001 to May 2026
  • The complete main site, browsable offline exactly as it appeared: 7,872 articles, 1,526 news items, and 153 featured-photo galleries
  • Reader comments (5,736 across 8,686 discussions) reunited with the articles they were written on
  • All 307 Wetpixel Live episodes saved as audio and video, with written transcripts of every show: over 80 hours of conversation
  • Images that had already vanished from the live site recovered from old backups and restored into the pages that reference them
  • The Wetpixel Wiki: an AI-written history of the site and the first 20 years of digital underwater photography, with nearly 300 pages on the people, gear, techniques, and events of the era, plus search that understands meaning as well as keywords
  • The entire record fits on a single drive (158 GB) and needs no server to browse

Under the hood

858Files
11,919Lines of code
52Commits
Jun 2026Dev window
StackPython 3 · wget · HTTrack · SQLite · yt-dlp · WhisperX · Disqus XML export · Astro · Cloudflare Pages · Workers AI · Vectorize
InfraLocal Mac / NAS for the archive; wiki published via Cloudflare Pages
Notable
  • Crawl-once / derive-everything architecture: a single wget master mirror is the source of truth; the HTTrack browse copy, SQLite dataset, and comment join are all produced offline, so the live server is hit exactly once and every stage is idempotent
  • Authenticated IPS4 forum crawl captures 412,725 posts across 64,016 threads from 14,960 members (December 2001 to May 2026)
  • Disqus XML export parsed and joined to archived content items by URL: 5,736 comments across 8,686 threads, with ~994 junk-host threads filtered at join time
  • WhisperX transcription of 307 Wetpixel Live episodes (80 hours 40 min); transcription packages distributed to remote machines for parallel processing
  • Cold-storage image recovery merges ~1.08 GB of legacy image directories into the browse copy and rewrites saved 404 stubs to point at recovered files
  • Wiki pages are LLM-generated from the enriched archive (articles, curated forum threads, news, and episode transcripts), organized as a year-by-year timeline plus 107 people profiles, gear, techniques, events, companies, and locations
  • Wiki ships as an Astro static site on Cloudflare Pages with dual search: keyword via Pagefind and semantic via Workers AI embeddings in Vectorize