Internet archive under siege as cyberattacks and rights demands mount

The Internet Archive, the world's largest digital repository, is facing a mounting crisis due to relentless cyberattacks and intellectual property claims. The non-profit, launched 30 years ago to preserve the digital landscape, now stores over a billion archived web pages, serving as a vital resource for journalists, researchers, historians, and legal experts.

Price of progress: the internet archive

Price of progress: the internet archive's struggle

Ranging from crippling cyber assaults to rights infringement lawsuits, the Archive faces an existential threat. In May 2024, the platform entirely collapsed for three days, highlighting the severity of the situation. Beyond security breaches, the Archive has deleted over 500,000 books due to questionable authorship.

The final blow may come from the very media outlets that rely on the Archive for their own historical record-keeping. An increasing number of publications, including The Guardian, The New York Times, and USA Today Co, have begun blocking their content from being stored on the platform. Some have explicitly stated their reluctance to have their work used to train AI models, while giants like OpenAI and Google are voraciously scraping the Archive for their own purposes.

As bots launched by these companies flood the Archive's servers with tens of thousands of requests per second, the repository's infrastructure buckles under the strain. Mark Graham, director of the Wayback Machine, confirms that some companies are accessing the archives on an unprecedented scale, hammering the system with automated queries.

Over 100 journalists have rallied in support of the Archive, recognizing the critical role it plays in preserving the web's historical record.