Extract URLs

Paste text, HTML source or an XML sitemap and get every web address in it as a clean list, one per line.

Advertisement

Uses

  • SEO audits: list every page in a sitemap, or every link on a page from its HTML source
  • Research: collect the links from notes, chats or documents
  • Migrations: get a list of URLs to check or redirect

Only addresses starting with http:// or https:// are found. Punctuation at the end of a sentence is trimmed off, and duplicates are removed.

Then alphabetize them, compare two sitemaps to find missing pages, or count how often each link appears.

From a sitemap

Before
<urlset>
  <url><loc>https://example.com/</loc></url>
  <url><loc>https://example.com/about/</loc></url>
  <url><loc>https://example.com/contact/</loc></url>
</urlset>
After
https://example.com/
https://example.com/about/
https://example.com/contact/

From a paragraph of text

Before
See https://example.com/docs, or the mirror (http://mirror.example.net/docs). Again: https://example.com/docs.
After
https://example.com/docs
http://mirror.example.net/docs

Questions

Does it find relative links like /about?

No. Only full addresses beginning with http:// or https:// are extracted.

Can I extract links from a PDF?

Copy the PDF’s text and paste it here. Links that show only as underlined words, without the address in the text, can’t be found this way.