Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.
  • Torches and pitchforks?
  • by the enclosure "movement"? i.e. greedy powerful people walling off everything they could and declaring it was theirs and you'd have to pay a tithe to use it?

    (I should really start calling rent "tithes" more often)

  • I recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.

    I really so through some money their way, they do wonderful work.

  • It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

    I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

  • I'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start:

    - charging (news / journalist services)

    - gate-keeping (X forcing log-ins)

    - enshittifying (lots of ads and degraded service)

    The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.

  • I've been doing a bit of (very careful to be polite) scraping of Wayback to get archives of now-offline sites, so I hope this won't cause any significant issues for me. I did contact them beforehand to request a direct copy of the sites/networks required (which I believe they offered at some point) but unfortunately received no reply.
  • > We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
  • I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

    Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

  • I wonder if your browser is prefetching every link you move the mouse over.
  • Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible.

    Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible to tell a real user from a bot. Only solution is to pretend that everyone is a bot and design for it.

  • Wow, but I wonder if there's more to it.

    I've not been able to access web.archive.org from my work computer - I always get the 429 error.

    But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

  • They seem to be aggressively blocking IPv6 source IPs. I ran into this problem over the past month traveling. I got nothing but 429 errors until I switched on my VPN (which is IPv4 only) and magically the Wayback machine worked again. The lack of transparency on the part of IA is very frustrating.
  • Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
  • My buddy said he could not access it even from a residential IP, it was blacklisted for some reason.
  • Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

    Try making a vpn via digital ocean for example and you'll see similar patterns.

  • I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
  • Weird that a library is restricting free access to other libraries wanting to preserve history. I guess it’s not a library after all.
  • The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic

    Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.

    I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.

  • Unrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point.

    Still can't remember what my Tripod site address was, but that might be lost to time.

    Thank you, Archive.org.

  • Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

    The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

    If you got some money to spare, consider donating to them. They need it.

  • fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

    The future is bleak :\

  • I've been donating $5 to them monthly for I don't know how long. I've only recently bumped it up to $25. They're the heroes of the internet age
  • I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.
  • > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

  • I think the Wayback Machine offers an official API for bots to call.
  • Nearly every time a link is posted to HN to a site behind some form of wall, a high voted comment on the post will be a link to an archive site bypassing the owners wall. Bot owners are not the only ones routinely circumventing the choices of content owners.
  • I wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point.

    Open access doesn't seem sustainable.

    But I might just grumpy about spending another hour this week adjusting rules to prevent bots.

  • Just yesterday from my one of my sessions with Sol:

    > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

  • I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)