A bulk Wayback Machine checker takes a list of URLs and asks the Internet Archive about every one of them at once, so you see in a single table which pages are archived and which are not. The free tool above checks up to 500 URLs per run, from a pasted list or a bookmarks file, and reports the first snapshot, last snapshot, the snapshot closest to any date you choose, and how many captures exist. For anything with no snapshot, it gives you a Save Page Now link so you can archive the page while it is still live.
Everything below is accurate as of September 2026.
What does a free bulk Wayback Machine checker do?
A bulk Wayback Machine checker runs the “is this archived?” question for many URLs at once instead of making you paste them one at a time into web.archive.org. It returns a yes or no per URL, plus the dates and links you need to open the right snapshot.
Typical bulk jobs:
- Check if URL is archived for every citation in a paper before you submit it.
- Find archived version of dead link entries in an old bookmarks folder.
- Audit a website’s outbound links before a migration.
The checker on this page returns, for each URL:
| Column | What it tells you |
|---|---|
| Archived | Yes if the Wayback Machine holds at least one successful capture, No if it holds none, Unknown if the lookup failed |
| Snapshots | How many successful (HTTP 200) captures the lookup found |
| First | Date of the oldest capture, linked to that snapshot |
| Last | Date of the newest capture, linked to that snapshot |
| Closest | The capture nearest your target date, and how many days away it is |
| Links | All captures (the calendar view), Near save date (for bookmarks with a known save date), or Save Page Now for unarchived URLs |
You can filter to show only archived, only unarchived, or only the rows that could not be checked, list unarchived URLs first, and download everything as a CSV file that opens in Excel, Numbers or Google Sheets.
How does the Wayback Machine work?
The Wayback Machine is the public front end to the Internet Archive’s web collection: billions of saved copies of web pages, each stored with the exact time it was captured.
The Internet Archive, a nonprofit library based in San Francisco, started crawling the web in 1996 and opened the Wayback Machine to the public in 2001. In October 2025 it marked one trillion archived web pages. Captures come from three main sources:
- Automated crawls. The archive’s own crawlers, plus crawls run with partner organisations, visit large parts of the public web on a rolling basis. Popular pages get captured often; obscure pages rarely or never.
- Save Page Now. Anyone can ask the archive to capture a specific live page at web.archive.org/save.
- Partner collections. Libraries, universities and governments run collecting projects through the Internet Archive’s Archive-It service, and some of that material is visible in the Wayback Machine.
Each capture is stored as a record in a WARC file (the standard web archive format) and indexed by URL and time. When you open a snapshot, the Wayback Machine replays the stored HTML and rewrites its links so images, stylesheets and other pages load from the archive too, ideally from captures made around the same moment.
How Wayback Machine URLs are built
Every snapshot has a predictable address, which is what makes bulk checking and citing possible. The timestamp is always UTC.
| URL pattern | What it opens |
|---|---|
web.archive.org/web/20150312093000/https://example.com/ |
The capture made at 09:30:00 UTC on 12 March 2015 |
web.archive.org/web/2015/https://example.com/ |
The capture nearest to the start of 2015 (partial timestamps redirect to the closest capture) |
web.archive.org/web/*/https://example.com/ |
The calendar of all captures of that URL |
web.archive.org/web/*/example.com/* |
A list of every archived URL under that domain |
web.archive.org/web/2015id_/https://example.com/ |
The original page as captured, without the Wayback toolbar or link rewriting |
web.archive.org/save/https://example.com/ |
The Save Page Now form for that URL |
If you saved a bookmark on 4 June 2019, web.archive.org/web/20190604/ followed by the URL takes you straight to the capture closest to that day. The checker’s “Near save date” link builds exactly this address from the date stored in your bookmarks file.
How do you check if a URL is archived?
To check if a URL is archived, open web.archive.org/web/*/ followed by the URL, or paste the URL into the search box on web.archive.org. A calendar with highlighted dates means it is archived; “Wayback Machine has not archived that URL” means it is not. For more than a handful of URLs, use the bulk checker, which does the same lookup for each one.
Behind that simple answer there are two official lookup APIs, and knowing how they differ explains why a bulk Wayback Machine checker is built the way it is. Both are documented on the Internet Archive’s Wayback Machine APIs page.
The Wayback Availability API
The Availability API answers “is there a snapshot, and which one is closest to this time?” You call archive.org/wayback/available?url=example.com×tamp=20150101 and get back JSON with a single closest object containing available, url, timestamp and status. With no timestamp, it returns the most recent capture.
It is fast, but it only ever returns one snapshot: no first capture date, no capture count, no coverage over time. It has also been known to return an empty result for URLs that do have captures when the archive is under heavy load, so a “no” from the Availability API alone is weak evidence.
Wayback CDX API basics
The CDX API is the archive’s capture index. It returns one row per capture of a URL, which makes it the right source for first and last dates and for counting captures. The full reference lives in the Wayback CDX server README.
A basic query looks like web.archive.org/cdx/search/cdx?url=example.com&output=json. Each row carries the fields urlkey, timestamp, original, mimetype, statuscode, digest and length. The parameters that matter most:
| Parameter | Example | What it does |
|---|---|---|
url |
url=example.com/page |
The URL to look up (scheme optional) |
matchType |
matchType=prefix |
exact (default), prefix (everything under a path), host, or domain (including subdomains) |
from / to |
from=2015&to=2018 |
Limit to a date range; accepts 4 to 14 digit timestamps |
limit |
limit=1 or limit=-1 |
First N rows, or with a negative number the last N rows |
filter |
filter=statuscode:200 |
Keep only rows matching a field; prefix with ! to exclude |
collapse |
collapse=digest |
Merge adjacent rows with the same value, e.g. identical page content |
fl |
fl=timestamp,original |
Choose which fields to return |
output |
output=json |
JSON rows instead of space-separated text |
A Wayback Machine bulk lookup done properly combines these: limit=1 with filter=statuscode:200 for the first good capture, limit=-1 for the last, and a closest-to-date query for the snapshot nearest your target. That is the approach the tool on this page uses through the Bookmend tools API, which also caches answers for 24 hours so repeat checks do not hit the archive again.
Why the checker filters to status 200
Filtering to HTTP 200 captures means “archived” really means “there is a copy of the page’s content.” Without the filter, a URL whose only captures are 404 errors or redirects to a homepage would show as archived, and opening it would show you an error page from the past. If a URL shows No but you believe it existed, the archive may only hold redirect captures for it; open the All captures calendar to see those too.
Wayback Machine bulk lookup: why 500 URLs, in batches of 10?
The checker caps each run at 500 URLs and sends them in batches of 10 because the Internet Archive is a free, donor-funded service with rate limits, and hammering it gets everyone blocked. A run of 500 URLs finishes in a few minutes.
- Duplicates are skipped by default.
http://example.com/page/andhttps://www.example.com/pagecount as one URL, so a messy list does not waste your 500. - Bookmarks files can be filtered by folder. If your export holds 4,000 links, pick the folder you care about. Subfolders are included.
- Progress is kept. If the archive rate-limits the lookup, the tool waits and retries automatically, and if that still fails you get a Resume button that continues from the first unchecked batch.
- Non-web links are dropped.
javascript:bookmarklets,chrome://pages,file://paths and Firefoxplace:queries cannot be archived, so they are skipped rather than reported as missing.
For tens of thousands of URLs, call the CDX API from a script with a sensible delay instead.
How do you read the results?
Read Archived first, then the dates. An archived URL with a recent Last date is well covered; one whose only capture is years old may be missing the version you want; an unarchived URL that is still live should be saved now.
Archived: yes, but check the dates
A Yes means at least one good capture exists. It does not mean every version of the page is saved, or that the capture is complete. JavaScript-heavy pages can replay as an empty frame, images may come from captures months apart, and a capture might be a cookie wall or “please enable JavaScript” screen. Open the snapshot before you rely on it.
Archived: no
A No means the lookup found no successful capture for that exact URL. Before giving up, try the variations in the troubleshooting table further down. Tracking parameters are the most common culprit: example.com/post?utm_source=newsletter is a different URL to the archive than example.com/post. Strip them first; the free URL cleaner removes utm, fbclid, gclid and similar parameters from a whole list.
Couldn’t check
An Unknown row means the lookup itself failed: a timeout, an archive outage, or a URL the API refused. It says nothing about whether the page is archived. Press Resume, or open the All captures link to check that URL by hand.
Closest snapshot and the target date
Set a target date when the version matters, such as the day you bookmarked a page. The Closest column then shows the nearest capture and how many days away it is. A gap of a few days is usually fine; a gap of years means the snapshot may show a different version of the page entirely.
How do you find an archived version of a dead link?
To find an archived version of a dead link, look it up in the Wayback Machine with a target date near when you last saw the page working, then open the closest snapshot. If that fails, try URL variations and a second archive such as archive.today.
A workflow that recovers the most links:
- Confirm the link is actually dead. A page that returns 403 or a Cloudflare challenge to a checker may load fine in your browser. The free bookmark dead link checker sorts a bookmarks file into dead, parked, redirected and needs-a-look, so you only chase the real losses.
- Run the dead URLs through this checker with a target date. For bookmarks, the Near save date link uses each bookmark’s own save date, which is more precise than one date for the whole list.
- Open the closest snapshot and check it. Make sure it shows the content, not an error page or a login prompt.
- Try variations for anything not archived. Remove the query string, switch
www.on or off, remove a trailingindex.html, or check the parent path withweb.archive.org/web/*/example.com/blog/*. - Check archive.today. Search the URL at archive.ph. Archive.today holds on-demand captures made by individuals, often of news articles and social posts.
- Replace the bookmark or citation with the archived URL. In Chrome, right-click the bookmark, choose Edit, and paste the snapshot address in the URL field.
Bookmend automates much of this for the bookmarks you save in it: its link health feature checks each saved link in the background (at most once a day per bookmark), flags links that have died, and offers the Wayback Machine snapshot closest to when the page last worked as a replacement. Bookmend also keeps the readable text of each page you save, so the content stays available even if the site changes. The post on how Bookmend recovers dead bookmarks from the Internet Archive walks through the recovery idea in more detail.
What does the Wayback Machine not archive?
The Wayback Machine does not archive pages that require a login, pages whose owners have asked to be excluded, most content assembled by JavaScript after the page loads, and pages no crawler or person ever pointed it at. The table covers the common gaps.
| What is missing | Why | What you can do |
|---|---|---|
| Pages behind a login (webmail, members-only forums, private social posts) | Crawlers and Save Page Now visit as an anonymous visitor | Nothing in public archives; keep your own copy (print to PDF) |
| Sites excluded at the owner’s request | The Internet Archive honours exclusion requests from site owners and shows “This URL has been excluded from the Wayback Machine” | Try archive.today or national web archives, which have different policies |
| Pages that block crawlers | Save Page Now will not capture pages that do not allow crawling | Try archive.today; save your own copy |
| JavaScript-heavy apps (single-page apps, infinite scroll, embedded maps) | The capture stores what the server sent; content fetched later by scripts may not replay | Look for a print view or an AMP/text version of the URL |
| Very new or obscure pages | Crawls prioritise pages that are linked from many places | Use Save Page Now while the page is live |
| URLs with session IDs or tracking parameters | Each unique URL is a separate entry in the index | Look up the clean URL instead |
| Large files and streaming media | Many crawls cap file size and skip streams | Check the Internet Archive’s other collections, or the original host |
The archive’s index also normalises some differences for you. In most cases http:// versus https:// and www. versus no www. point at the same set of captures, which is why the checker treats those as duplicates.
What should you do when a URL is not archived?
If a URL is not archived and the page is still live, click Save Page Now next to it and the Internet Archive will capture it within a minute or so. If the page is already gone, Save Page Now cannot help; look in other archives instead.
Save Page Now is free and open to everyone. Logging in to a free archive.org account adds extra options, such as capturing the page’s outlinks and seeing a list of your own captures. You can also email a list of links to the archive’s Save Page Now address and get a status reply, which suits longer lists.
The checker deliberately does not submit pages for you. Archiving makes a public, permanent copy, and some pages (a private document shared by link, a page with your account details in the URL) should never be archived. Clicking each Save Page Now link keeps that decision with you.
When a save fails, the usual reasons are that the site blocks the archive’s crawler, the page needs a login, the site is slow and the capture timed out, or you have hit the rate limit for anonymous saves. Wait a few minutes and try again, or try archive.today, which works differently.
Wayback Machine checker alternatives: which archived website finder should you use?
The Wayback Machine has the largest collection, so check it first; a bulk Wayback Machine checker is the fastest way to do that for a list. Other tools are better for single URLs, for pages the Internet Archive lacks, or for keeping your own bookmarks alive without manual checks.
- Bookmend Our product — an AI bookmark manager for Chrome and the web. Import your bookmarks.html once and its link health feature keeps checking the saved links in the background, flags dead ones, and offers the Wayback Machine snapshot closest to when each page last worked, with no list to paste. Link health is a Premium feature and free during early access.
- This bulk checker — best when you have a list of URLs (citations, a site’s outbound links, an exported bookmarks folder) and want archived yes or no plus dates for all of them at once, free, in a browser.
- The official Wayback Machine browser extension — from the Internet Archive; offers the archived copy when you hit a missing page and saves the page you are on with one click. One page at a time.
- web.archive.org directly — the calendar view is still the best way to browse every version of one page and compare changes over time.
- Archive.today (archive.ph, archive.is) — an independent on-demand archive that keeps a screenshot with every capture and sometimes has pages the Wayback Machine does not. There is no official bulk lookup. Our guide to archive.today covers how it differs.
- Memento Time Travel — an aggregator built on the Memento protocol (RFC 7089) that searches several public web archives for one URL at a time.
- Scripts against the CDX API — the right choice for tens of thousands of URLs, if you are comfortable with code and rate limits.
For a longer comparison of archives and when each one wins, see Wayback Machine alternatives.
| Tool | URLs per lookup | First/last dates | Closest to a date | Cost |
|---|---|---|---|---|
| Bookmend (AI bookmark manager) | Every bookmark saved in it, in the background | Not shown; offers the closest | Yes, closest to when the page last worked | Premium (free in early access) |
| This bulk checker | Up to 500 per run | Yes | Yes, any date you choose | Free |
| web.archive.org calendar | 1 | Yes, visually | Yes, by clicking | Free |
| Wayback Machine extension | 1 (the current page) | Via links to the calendar | Via the calendar | Free |
| Archive.today | 1 | Lists its own captures | No | Free |
| CDX API scripts | Unlimited, rate-limited | Yes | Yes | Free, needs code |
How do you cite an archived web page?
Cite an archived page by giving the original title and URL as usual, then adding the full Wayback snapshot address and the capture date. The 14-digit timestamp in the snapshot address pins the exact version you read.
- Academic styles (APA, MLA, Chicago). Cite the page normally and use the archived URL, or add it after the live URL, so a reader can reach the version you saw if the original changes or disappears.
- Wikipedia. The
{{cite web}}template hasarchive-urlandarchive-datefields, and a bot called InternetArchiveBot adds archive links to dead citations automatically. - Legal and journalistic work. Save the page yourself at the moment you cite it rather than relying on an existing capture, so the timestamp matches your access date.
The CSV export from this tool is useful here: run every URL in a reference list, filter to Not archived, save the live ones, then rerun and paste the closest snapshot addresses into your bibliography. If you also need to know when the page itself was last changed, the page’s visible dates and HTTP headers are separate from the Wayback capture date.
Troubleshooting: why does the checker show something unexpected?
Most surprises come from URL variations, redirect-only captures, or archive outages.
| Symptom | Likely cause | Fix |
|---|---|---|
| “No” for a page you know was popular | The exact URL differs from the archived one (tracking parameters, mobile m. subdomain, trailing index.html) |
Look up the clean URL, or open web.archive.org/web/*/domain.com/* to find the archived form |
| “No” but the calendar shows dates | Only redirect or error captures exist; the checker counts HTTP 200 captures | Open All captures; blue dates are good captures, green are redirects |
| Snapshot opens a blank or broken page | Content was loaded by JavaScript, or assets were not captured | Try an earlier or later capture, or the id_ raw view |
| Snapshot shows a cookie wall or login page | That is what the crawler saw at capture time | Try other captures around that date |
| “Couldn’t check” on many rows | The Internet Archive is slow, down, or rate-limiting | Wait a few minutes and press Resume |
| Run button is disabled | Live checking is being switched on for this tool, or a run is already in progress | Check back shortly, or wait for the current run |
| Fewer URLs checked than pasted | Duplicates removed, non-web links dropped, or more than 500 URLs | See the count line under the button; split long lists |
| Bookmarks file not recognised | The file is not a bookmarks export | Export bookmarks.html from your browser (Chrome: Bookmark manager, ⋮ menu, Export bookmarks) |
| Save Page Now fails | Site blocks crawlers, needs login, timed out, or anonymous save limit reached | Retry later, log in to archive.org, or try archive.today |
Is this Wayback Machine checker free and private?
Yes. The checker is free, needs no account, and only sends the URLs you are checking to the Bookmend tools API, which forwards the lookup to the Internet Archive. A bookmarks file is parsed in your browser tab and never uploaded; titles, folders and save dates stay on your device.
Do not paste URLs that contain private tokens (password-reset links, shared-document links with access keys), and remember that anything you save with Save Page Now becomes public.
Keep your own bookmarks recoverable
A one-off bulk check tells you which links can be recovered today. Bookmarks keep dying after that, and the longer a page has been gone, the harder it is to find the version you saved.
Bookmend handles the ongoing part. It is an AI-native bookmark manager that works as a Chrome extension (also in Brave, Edge, Arc and Opera) and as a web app. You import a bookmarks.html file from Chrome or any other browser, up to 20,000 bookmarks at a time with their folders, and the library lives in your Bookmend account. Its link health feature then checks your saved links in the background, at most once a day per bookmark, flags the dead ones and offers the Wayback Machine snapshot closest to when each page last worked as a replacement. Because Bookmend keeps the readable text of every page you save, you can also search your library by meaning or ask it questions in plain language and get answers with citations back to your bookmarks.
The free plan covers the bookmark manager itself: saving, collections, tags, search, and import and export. Link health and the AI features are part of Premium, and during early access they are free for every new account, with no card needed. Note that Bookmend’s servers fetch the pages you save in order to read them, and AI features send page text to model providers. See the features or how it works for the details.
Written by Nafiul Hasan, Founder, Bookmend. Last updated.