You can use the Wayback Machine to rebuild an expired domain's site safely if you treat the archive as a map, not a source of content. Recover the old URL structure and topic focus, prioritize the URLs that still have backlinks, and publish new, original content on the same subjects at those URLs. Do not republish the archived text wholesale. It is still copyrighted by its original author, and a copy-paste revival is exactly the pattern Google's "expired domain abuse" spam policy targets. This guide covers when a rebuild makes sense, the legal and policy lines, and a practical workflow using the Wayback CDX API.
Safe Rebuild Workflow
When Restoring an Old Site Makes Sense
A rebuild is not the default for every expired domain. It makes the most sense when the domain's past and your plans line up:
Good candidates
- Clear, consistent topical history (for example, years as a hobby or how-to site)
- Backlinks pointing at specific inner pages, not just the homepage
- A topic you can genuinely cover with expertise
- Clean history: no spam, hacks, or parked-ad periods (see Spam Checking)
- A simple URL structure that's easy to recreate
Poor candidates
- Former government, school, charity, or medical sites you would turn commercial
- Topic you have no plans or competence to cover
- History dominated by thin, scraped, or spun content
- A well-known brand whose audience would be misled by a revival
- Value comes only from homepage links (a 301 redirect or fresh money site may fit better)
If the old site was a recognizable organization or person, be careful even with a topic match. Presenting yourself as the continuation of a defunct business or author is misleading to users and can raise trademark and passing-off issues. See Trademark & Legal Risk.
Copyright: Archived Content Isn't Yours
Buying a domain name transfers control of the name, nothing else. The text, images, photos, and code on the old site are still owned by whoever created them (the former owner, their writers, photographers, or licensors), whether or not that person still exists as a business. The Internet Archive keeps copies for preservation and research. That does not give anyone a license to republish them.
- Republishing archived articles verbatim is copying protected work. The rights holder can send DMCA takedown notices to your host and to Google, which can remove the pages from search.
- Images are the highest-risk item. Stock photos and photographers' work are actively monitored by rights-enforcement services, and the original site's license (if any) doesn't carry over to you.
- Facts and ideas aren't copyrightable; expression is. You can cover the same topics, answer the same questions, and use the same URL slugs. You need to write it yourself.
- Permission is an option. If the former owner is reachable, a written license to reuse specific content can be negotiated. Get it in writing and keep it.
Not legal advice: Copyright law varies by country, and fair use and fair dealing are narrow, fact-specific defenses that do not cover republishing whole articles. If you plan to reuse anything substantial, consult a lawyer.
Google's Expired Domain Abuse Policy
In March 2024, Google expanded its spam policies with three new categories: expired domain abuse, scaled content abuse, and site reputation abuse. These are the policies that matter most for archive rebuilds.
What Google Considers Abuse
Google defines expired domain abuse as buying an expired domain and repurposing it primarily to manipulate search rankings by hosting content that offers little or no value to users. Google's own examples follow a pattern: the new content trades on a reputation built for a completely different purpose. Examples include affiliate content on a former government site, commercial medical products on a former non-profit health charity's domain, and casino content on a former school's site.
What Is Legitimate Reuse
Google is explicit that buying a previously used domain is not itself a problem. Legitimate reuse typically looks like:
- A new site on the same or a closely related topic, built to serve the audience the old links and visitors expect
- Original, helpful content created for users, not just to capture the old rankings
- Honest presentation: no pretending to be the old organization, no fake "since 2009" claims
- A brand-new project that simply likes the name, without trying to exploit the old ranking signals
Related Policies to Avoid Tripping
- Scaled content abuse: mass-producing pages (by AI, scraping, spinning, or templates) mainly to rank. Restoring hundreds of URLs with auto-generated text falls here, even if the topics match.
- Site reputation abuse: third-party content published on a host site to exploit that site's ranking signals, with little oversight from the host. Renting out sections of a revived domain to unrelated publishers risks this.
Enforcement: Google can act on these policies algorithmically or through manual actions. A clean Search Console report on day one doesn't mean the site will stay clean. What you publish afterwards is what gets judged. See Google Index & Penalty Check.
Recovering the URL Structure with the CDX API
The Wayback Machine's web interface is fine for reading individual snapshots, but listing every URL it captured is easier with the CDX Server API. It returns one row per capture, with filtering and de-duplication.
Basic Query
This request lists every unique URL captured under a domain, with the original URL and HTTP status:
http://web.archive.org/cdx/search/cdx?url=example.com/*&output=json&fl=original,statuscode&collapse=urlkey
url=example.com/*: the trailing wildcard matches every path under the host (the same asmatchType=prefix). UsematchType=domainto include subdomains.output=json: returns a JSON array (the first row is the field header). Omit it for space-separated text.fl=original,statuscode: limits output to the fields you need. Other useful fields aretimestamp,mimetype, anddigest.collapse=urlkey: collapses repeated captures of the same normalized URL into one row.
Useful Refinements
- Only successful pages: add
&filter=statuscode:200to drop redirects and errors. - Only HTML: add
&filter=mimetype:text/htmlto skip images, CSS, and scripts. - Limit to the "good" era: use
&from=2015&to=2019(timestamps can be shortened to years) to exclude spam or parked periods you found during vetting. - Large sites: use
&limit=with&showResumeKey=trueto page through results instead of making one huge request.
To see a capture's raw HTML without the Wayback toolbar or rewritten links, add id_ after the timestamp: https://web.archive.org/web/20180101000000id_/http://example.com/page. That makes it easier to read the old headings, titles, and internal link structure.
Be a polite client: the Internet Archive is a non-profit, and it rate-limits heavy traffic. Space out requests, cache responses locally, and avoid downloading far more than you need.
Prioritize URLs That Have Backlinks
A CDX export often lists thousands of URLs, most of them tag pages, pagination, feeds, and query strings. You don't need to rebuild them all. Join the archive list with your backlink data:
- Export linked pages from Ahrefs (Best by links) and Majestic (Pages), sorted by referring domains.
- Match them against the CDX list (normalize trailing slashes,
www, and protocol first). - Sort into tiers:
- Tier 1: URLs with several quality referring domains. Rebuild these at the exact same path.
- Tier 2: URLs with a few links, or links from weaker sites. Rebuild if the topic fits, otherwise 301 to the closest new page.
- Tier 3: no links. Ignore, or rebuild only if the topic belongs in your new site plan anyway.
- Check each Tier 1 snapshot to note the original title, headings, and purpose of the page so your replacement matches its intent.
The backlink and Wayback vetting methods are covered in Backlink Analysis and History & Wayback Machine.
Rebuild with New Original Content
The goal is pages that a reader following an old link would find at least as useful as the original, written by you.
Matching the Original Topic
- Keep the URL where possible. That's what the backlinks point to.
- Keep the subject and intent: if the old page was a beginner's guide to repotting orchids, the new one should be too, not a product roundup.
- Write from scratch: use the archived page only to understand what it covered. Don't paraphrase it sentence by sentence; that can still be a derivative copy.
- Update and improve: current information, better structure, your own photos or diagrams, clear authorship.
- Be transparent: an "About" page saying the site is under new ownership is honest and costs you nothing.
Illustrative Example
Hypothetical scenario: example-gardening.com was a hobby blog from 2012–2019. CDX shows about 400 HTML URLs, and backlink data shows that 25 of them hold most of the referring domains, mostly care guides for specific houseplants. A safe rebuild publishes 25 newly written care guides at those exact paths, adds new related content over time, 301s a handful of weakly linked posts into the closest new guide, and lets the rest 404. An unsafe rebuild downloads all 400 pages with a scraper and republishes them, or swaps the topic to unrelated affiliate reviews.
Tools for Archive-Based Rebuilds
- Wayback Machine (web.archive.org): browse snapshots and the calendar view to identify good eras. Free.
- CDX Server API: bulk URL discovery, as shown above. Free.
- Open-source Wayback downloaders: several community projects (in Ruby, Python, and Go, for example) fetch archived snapshots in bulk. They are useful for offline research, such as mapping site structure, internal links, and page titles. Using them to republish the downloaded content is where the copyright and spam risks come in.
- Backlink tools: Ahrefs, Majestic, Semrush for prioritizing URLs.
- Spreadsheet or script: to join CDX output with backlink exports and track an action per URL.
- Crawler (e.g. Screaming Frog): after launch, to confirm every Tier 1 and Tier 2 URL returns 200 or a single 301.
Commercial archive-restoration services also exist. If you use one, you are still responsible for the copyright status of whatever it restores.
Pre-Launch Checks for a Rebuilt Site
Before you point search engines at the rebuilt site, run through a short quality gate:
- Originality: spot-check rebuilt pages against their archived versions. No copied sentences or lightly reworded paragraphs.
- Status codes: every Tier 1 URL returns 200, every redirected URL returns a single 301, and spam-era URLs return 410.
- Internal linking: restored pages link to each other the way a real site would, not only from a sitemap.
- Authorship and ownership: real author information and an honest About page are in place.
- Media rights: every image is your own, properly licensed, or public domain.
- Outbound links: old affiliate or sponsored links are removed, or marked
rel="sponsored"if they are intentional and current.
What Not to Do
- Don't republish archived content wholesale, including via "rewriter" or spinning tools that produce near-copies.
- Don't reuse old images unless you have confirmed rights to them.
- Don't impersonate the former owner: no copying their logo, author bios, testimonials, or "about us" story.
- Don't pivot to an unrelated, monetization-first topic while relying on the old reputation. That is the core of expired domain abuse.
- Don't mass-generate pages to fill every archived URL. Hundreds of thin pages invite scaled content abuse issues.
- Don't restore the spam era: if the domain was hacked or hosted spam at some point, exclude those URLs (return 410) rather than rebuilding them.
- Don't restore old outbound links blindly. Old sponsored or paid links in archived posts shouldn't reappear as followed links on your site.
Once the rebuilt pages are live, follow the post-acquisition checklist for Search Console verification, sitemaps, and monitoring.
FAQ
Is it legal to republish content from the Wayback Machine on my expired domain?
Generally no, not without permission. Buying the domain doesn't transfer copyright in the old site's text and images, and the Internet Archive's copies don't grant a license. Use archived pages for research, then write new original content.
Does Google penalize rebuilding an expired domain's old website?
Not by itself. Google's expired domain abuse policy targets domains repurposed mainly to manipulate rankings with low-value content, especially content unrelated to the domain's past. A new, original site on the same topic that serves users is legitimate reuse.
How do I get a list of every URL an old site had?
Query the Wayback CDX API, for example web.archive.org/cdx/search/cdx?url=example.com/*&output=json&fl=original,statuscode&collapse=urlkey. Add filter=statuscode:200 and filter=mimetype:text/html to keep only real pages.
Which archived URLs should I rebuild first?
Prioritize the URLs with the most quality referring domains in Ahrefs or Majestic. Rebuild those at the same path with new content on the same topic, 301 weaker ones to close equivalents, and let unlinked URLs 404.
Can I use a Wayback Machine downloader tool?
Open-source downloaders are fine for offline research, such as mapping structure, titles, and internal links. Republishing what they download carries the same copyright and spam-policy risks as copying by hand.
Next Steps
Continue with these related guides:
- History & Wayback Machine — Audit a domain's past before you buy
- Post-Acquisition Checklist — The first 30 days after buying an expired domain
- Trademark & Legal Risk — Avoid brand and copyright problems
- Money Site Creation — Build a new site on an aged domain