discovered 03 Aug 2026
domain-discovery-crawler
→ View on GitHubThe Domain Discovery Crawler is a Scrapy-based tool designed for large-scale, focused web crawling, utilizing Redis queues and a deep-learning-based link classification model. Its primary use case is to efficiently gather relevant domains based on seed URLs, while allowing customizable parameters for crawling behavior, domain relevance scoring, and integration with Redis for managing queues. Notable features include the ability to handle pagination, log queue statistics, support for autologin, and exporting crawl results in various formats.