discovered 03 Aug 2026
crawley
→ View on GitHubCrawley is a web crawling tool that efficiently parses HTML pages to discover and print links, including resources like images, audio, and video. It features a fast SAX-parser, customizable scan depth, compliance with `robots.txt`, support for subdomain crawling, and options for ignoring certain URLs or scanning specific tags. Additionally, Crawley allows for the use of user-defined cookies and headers, making it adaptable for various crawling scenarios.