But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
Two of my favorite sites [0][1] still online—Lurkers Guide to Babylon 5 and ex-astris-scientia—were started in ‘92 and ‘98 respectively.
Elegant content for a more civilized age.
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
0.mk, you had one job…
The plot features two American men who stumble upon Brigadoon, a mysterious Scottish village that appears for only one day every 100 years; one man soon falls in love with a young woman from Brigadoon. The show's song "Almost Like Being in Love" subsequently became a standard.That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.