Bookmend
← BlogSAVING & ARCHIVING THE WEB

How to Find an Archived Version of Any Website

By
⚡
THE SHORT ANSWER

Go to web.archive.org, paste the dead URL into the search box, and pick a snapshot date from the calendar that appears — it's free, no account needed. If the Wayback Machine doesn't have it, try archive.today (archive.ph) next, especially for paywalled news pages. Google's cache feature was removed in early 2024, so it's no longer an option. Bookmend automates this exact lookup for dead bookmarks, matching the closest snapshot to when you originally saved the page.

How to Find an Archived Version of Any Website

Pages disappear. Articles get deleted. Whole websites go dark overnight. When you need to find a version of a page that no longer exists — or one that existed at a specific point in time — there are reliable ways to do it. This guide walks through each method, when it works, when it doesn’t, and how to deal with the large category of bookmarks that quietly stopped working.

Why Pages Disappear (and Why It Matters)

Before getting into the tools, it’s worth understanding why archived pages are necessary in the first place.

The web is not stable. Studies of academic and news URLs consistently find that 20-30% of links in any large collection stop resolving within a few years. Individual link collections can have higher failure rates depending on what you saved. The causes:

  • Publication or business closures: A website goes offline entirely when the company stops paying for hosting or shuts down
  • CMS migrations: Sites move to new platforms and break URL structures, turning /blog/2019/article-title into /articles/article-title or something entirely different
  • Article deletions: Content gets removed for legal reasons, correction policies, or editorial decisions
  • Domain expiry: An owner doesn’t renew a domain, which is then purchased by a domain squatter showing ads or malware
  • Paywalls: Content that was free moves behind a paywall, so while the URL resolves, the content you originally read is no longer accessible
  • Soft-404s: Pages return HTTP 200 (success) but show generic error messages, “article not found” banners, or homepage content — the URL resolves but the content is gone

This last category is particularly tricky. A URL that returns 200 will look “alive” to any basic checker, but if it now shows a parking page or a “content not found” message, the article you wanted is gone. Good link-rot detection tools (including Bookmend) specifically check for soft-404 patterns, not just HTTP status codes.

The Wayback Machine: The Primary Archive Tool

The Internet Archive’s Wayback Machine (web.archive.org) is the largest public web archive in existence. It has been crawling and saving copies of web pages since 1996 and has archived well over a trillion web pages. It’s free, nonprofit, and the standard first tool for recovering a gone page.

How to Look Up a Specific Page

  1. Go to web.archive.org
  2. Paste the full URL of the page you’re looking for into the search bar
  3. Press Enter

If the page has been archived, you’ll see a calendar view with years across the top. Click a year to see the months and days for that year — colored circles indicate which days have captures.

What the circle colors mean:

  • Blue circle: Standard snapshot captured successfully
  • Green circle: Snapshot captured, no change detected from the previous capture
  • Orange/yellow circle: The crawl returned a redirect (3xx) for that URL on that date
  • Red circle: The crawl returned an error (4xx or 5xx) — the page was already dead at that point

Reading the Calendar to Find the Right Snapshot

Click a day with a blue or green circle to see timestamps. If multiple snapshots were taken that day, you’ll see a list of times. Click any timestamp to load the archived version as it appeared at that moment.

This matters when you want a version from a specific period. For example:

  • Finding an article as originally published before it was significantly edited
  • Seeing what a website looked like before a redesign
  • Recovering an article you read in a specific year and remember the approximate timeframe

Getting the Closest Snapshot to a Specific Date via API

If you know the approximate date you saw a page and want the nearest snapshot, the Wayback Machine has a simple API endpoint:

https://archive.org/wayback/available?url=YOURURL&timestamp=YYYYMMDD

Replace YOURURL with the page’s URL and YYYYMMDD with a date. It returns JSON containing the URL of the nearest available snapshot. For example, timestamp=20230315 finds the closest capture to March 15, 2023.

This is what Bookmend uses internally when recovering dead bookmarks — it queries the Wayback Machine for the snapshot closest to the date you originally saved the link, so you get the version you intended to read rather than an arbitrary later snapshot.

Requesting a New Save

If a page exists right now and you want to ensure it’s archived before it might disappear:

  1. Go to web.archive.org
  2. In the “Save Page Now” box on the right side of the homepage, paste the URL
  3. Click Save Page

The Wayback Machine immediately crawls and saves a snapshot. The process takes a few seconds to a minute. You’ll get a link to the archived version once complete. This is free and available without an account.

What the Wayback Machine Can and Cannot Preserve

The Wayback Machine captures HTML, images, CSS, and most static assets. It does not fully preserve:

  • JavaScript-heavy interactive elements: Dynamic content that requires running JavaScript may not render correctly in older snapshots. Newer captures have improved JavaScript rendering
  • Login-gated content: Pages that require you to be logged in to view the content weren’t accessible to the public crawler, so they typically aren’t archived
  • Paywall content: Content behind a paywall at crawl time is generally not captured
  • Video and audio: Embedded media from external platforms (YouTube, Vimeo, SoundCloud) is typically not preserved in the snapshot
  • Pages that blocked crawlers: Some sites use robots.txt to block the Wayback Machine’s crawler. Those pages may not be archived at all — or historical captures before the opt-out may still be accessible

Can websites opt out? Yes. Adding a Robots.txt Disallow for the Internet Archive’s user agent (ia_archiver) requests exclusion from new crawls. The Internet Archive generally honors these requests for new captures. However, snapshots captured before the opt-out request typically remain accessible.

What Happened to Google Cache

For years, Google cached a copy of most indexed pages, accessible via the cache: search operator or a “Cached” link in search results. This was often the fastest way to access a recently-deleted page.

Google deprecated and removed the web cache feature in early 2024. The cache: operator no longer works in Google Search, and cached links no longer appear in results. Google’s stated reason was that the Wayback Machine and other archiving tools serve this purpose more appropriately.

For practical purposes: Google cache is no longer an option. If your workflow relied on it, the Wayback Machine is now the replacement.

archive.today: A Second Independent Archive

archive.today (also accessible as archive.ph) is a separate archiving service with no affiliation to the Internet Archive. It operates differently from the Wayback Machine in a few important ways:

  • User-submitted snapshots: archive.today primarily archives pages that users explicitly request, rather than mass-crawling the web
  • Paywall article captures: archive.today is known for capturing paywalled news articles because it saves the rendered page as a visitor sees it — including pages that show content before a paywall prompt, or during a first-visit preview window
  • Permanent storage: Snapshots are intended to be permanent; the site does not honor robots.txt opt-out requests for already-submitted snapshots
  • No account required: Anyone can submit a save request or browse existing snapshots

To search archive.today: go to archive.today and paste a URL into the search box. If the page has been previously submitted, you’ll see a list of snapshots. If not, you can submit a new save request.

archive.today is particularly useful for:

  • Paywalled news articles from major publications
  • Pages on sites that block the Wayback Machine’s crawler
  • Controversial content that the original site may have been motivated to delete

Using Google’s Site: Operator to Confirm a Page Existed

If you’re trying to verify that a page existed (not access its content), Google’s site: operator can help even for deleted pages, briefly:

site:example.com/path/to/page

If the page was recently removed, Google’s index may still contain the indexed snippet for a short period before it’s de-indexed. You won’t get the content, but the title and description from the search snippet can help you confirm the page existed and what it covered — useful context when searching archives.

Common Crawl: For Researchers, Not Casual Users

Common Crawl is a publicly available web crawl dataset used primarily for machine learning research and large-scale web analysis. It contains petabytes of crawl data but isn’t designed for casual page lookup — there’s no simple search interface for finding a specific URL. Some third-party tools and academic projects are built on top of Common Crawl data. If you’re doing research at scale rather than recovering a specific page, Common Crawl is worth knowing exists.

Browser Extensions for Automatic Archiving

Several browser extensions aim to automatically archive pages as you browse:

  • Wayback Machine browser extension: Available for Chrome and Firefox, adds a button to check if the current page has been archived and submit a save request. When you hit a 404, it automatically checks if a Wayback Machine version is available and offers to redirect you.
  • SingleFile: Saves a complete offline copy of any web page as a single HTML file, including styles and images, stored locally. Useful for personal archiving of important pages.

The Wayback Machine extension’s automatic 404 interception is practical if you frequently encounter dead links while browsing — it intercepts the error page and shows whether an archive exists.

If you have a bookmark collection of any age, a portion of those bookmarks are already dead. The percentages vary, but a collection with bookmarks from 3-5 years ago typically has 15-30% dead links, and collections from 7+ years ago can be higher.

The Wayback Machine is the recovery tool for dead bookmarks, but manually looking up each broken link is tedious if you have hundreds or thousands of bookmarks.

Bookmend automates this process inside Chrome. It monitors your Chrome bookmarks, confirms each failure twice 48 hours apart before flagging a link dead (to avoid false positives from temporary outages), and then surfaces the closest Wayback Machine snapshot to the date you originally saved each link. You see a list of dead bookmarks with one-click access to the archive version — without manually doing Wayback searches for each one.

The free rot check tool at /rot-check on this site lets you paste in a bookmarks.html export and see which links are dead, entirely client-side with nothing uploaded. It’s a useful starting point to understand the scale of the problem in your specific bookmarks before deciding how to address it.

Archiving Your Own Website for Preservation

If you manage a website and want to ensure it gets archived before you shut it down or migrate it:

Submit to the Wayback Machine: The Wayback Machine’s “Save Page Now” tool can archive individual pages. For a full site, the Wayback Machine’s Save Page Now 2 (SPN2) API supports recursive crawls and can be used to systematically archive a site. Submitting a sitemap URL is often the most efficient approach for large sites.

Notify the Internet Archive directly: For significant or culturally relevant sites, you can contact the Internet Archive directly. They prioritize archiving websites in categories they believe are at risk of disappearing.

Generate a static mirror: Tools like wget --mirror or httrack can crawl your site and generate a static HTML copy that can be hosted or stored elsewhere. This is more complex but produces a fully self-contained archive independent of any third-party service.

archive.today submission: For smaller sites or specific important pages, submitting to archive.today provides a second independent archive with different coverage.

Finding Archives of a Page When You Don’t Know the Exact URL

Sometimes you know a page existed but can’t remember the exact URL. A few approaches:

Google search with site: operator: If you remember the domain, site:domain.com keyword may surface indexed snippets even for deleted pages (briefly after deletion).

Wayback Machine full-text search: The Wayback Machine has a full-text search feature (different from the URL lookup) that can search the text of archived pages. This is available at archive.org/search. Coverage isn’t complete but it’s useful for finding pages you remember reading but not the exact URL.

archive.today search: Similar URL-lookup capability, useful if the page might have been submitted there.

Google cached snippets (historical): This no longer works since Google removed caching in 2024.


If you have bookmarks pointing to pages that may have gone offline, run a free check with Bookmend’s rot check tool — no upload, no account, works locally in your browser. For a broader look at how link rot affects bookmark collections and what to do about it, see the Bookmend features page.

Frequently asked questions

Yes, completely. web.archive.org is free to browse, search, and submit save requests without an account. The Internet Archive is a 501(c)(3) nonprofit funded by donations and grants. You can use Save Page Now, browse the calendar view, and access any archived page at no cost. Creating an account is optional and enables features like bulk save requests.

The earliest web captures date to 1996, though coverage before 2000 is sparse and uneven. For most mainstream websites, captures going back 10-20 years are typically available. For niche or smaller sites, coverage may be limited to the past few years or nonexistent.

Several reasons a page might not be archived: the site blocked the crawler via robots.txt, the page was behind a login or paywall at crawl time, the site was obscure and not frequently crawled, or the page was added and removed between crawl cycles. In this case, check archive.today as a second option. If neither has it, the content may genuinely be unrecoverable through public archives.

Generally yes — the archived content is a copy of what was publicly available on the original site. For academic citations, the Wayback Machine URL with timestamp is a recognized citation format in many style guides. For legal or professional contexts, verify the requirements with whoever is receiving the citation.

The Wayback Machine browser extension has an Auto Save Page option that can archive pages you visit. For systematic archiving tied to bookmarks specifically, Bookmend monitors whether your bookmarked pages go dead and provides Wayback Machine recovery links when they do — combining monitoring with recovery in one workflow.

Two common reasons: the URL was reused for different content over time (the site published a new article at the same path), or the snapshot captured a redirect or error page rather than the content. Use the calendar to browse captures over time and look for the snapshot closest to when you originally saw the page. Multiple captures on the same day at different times sometimes show different versions.

Yes. Adding a Robots.txt Disallow entry for the Internet Archive's user agent (ia_archiver) requests exclusion from new crawls. The Internet Archive generally honors these requests for new captures. However, snapshots captured before the opt-out request typically remain accessible.

Google deprecated and removed the web cache feature in early 2024. The cache: search operator no longer works in Google Search, and cached links no longer appear in results. Google's stated reason was that the Wayback Machine and other archiving tools serve this purpose more appropriately. For practical purposes, Google cache is no longer an option.

archive.today (also accessible as archive.ph) is a separate archiving service that primarily archives pages users explicitly request rather than mass-crawling the web. It is particularly useful for paywalled news articles from major publications, pages on sites that block the Wayback Machine's crawler, and content that the original site may have been motivated to delete. It does not honor robots.txt opt-out requests for already-submitted snapshots.

Blue indicates a standard snapshot captured successfully. Green means a snapshot was captured with no change detected from the previous capture. Orange or yellow means the crawl returned a redirect for that URL on that date. Red means the crawl returned an error such as a 404 or 500 — the page was already dead or unavailable at that point.

Yes. The Wayback Machine has an API endpoint for this: https://archive.org/wayback/available?url=YOURURL&timestamp=YYYYMMDD — replace YOURURL with the page's URL and YYYYMMDD with a date. It returns JSON containing the URL of the nearest available snapshot. This is what Bookmend uses internally when recovering dead bookmarks, finding the snapshot closest to the date you originally saved each link.

Nafiul Hasan
WRITTEN BY

Nafiul Hasan is the founder of Bookmend. He builds tools that fix small, specific frustrations rather than reinvent them — Bookmend started after losing a decade of bookmarked research links to dead pages he couldn't get back. He writes here about link rot, how the Wayback Machine actually works, and what's real versus marketing in the browser-extension space.

Drafted with AI assistance, reviewed and edited by Nafiul Hasan.How we write

FounderChrome extension builderWrites about link rot & bookmarks