Uses
- SEO audits: list every page in a sitemap, or every link on a page from its HTML source
- Research: collect the links from notes, chats or documents
- Migrations: get a list of URLs to check or redirect
Only addresses starting with http:// or https:// are found. Punctuation at the end of a sentence is trimmed off, and duplicates are removed.
Then alphabetize them, compare two sitemaps to find missing pages, or count how often each link appears.
From a sitemap
<urlset> <url><loc>https://example.com/</loc></url> <url><loc>https://example.com/about/</loc></url> <url><loc>https://example.com/contact/</loc></url> </urlset>
https://example.com/ https://example.com/about/ https://example.com/contact/
From a paragraph of text
See https://example.com/docs, or the mirror (http://mirror.example.net/docs). Again: https://example.com/docs.
https://example.com/docs http://mirror.example.net/docs
Questions
Does it find relative links like /about?
No. Only full addresses beginning with http:// or https:// are extracted.
Can I extract links from a PDF?
Copy the PDF’s text and paste it here. Links that show only as underlined words, without the address in the text, can’t be found this way.