discovered 03 Aug 2026
browsertrix-crawler
→ View on GitHubBrowsertrix Crawler is a high-fidelity web crawling tool that operates within a Docker container and utilizes Puppeteer to control multiple Brave Browser instances concurrently. It captures web data via the Chrome Devtools Protocol, enabling configurable and intricate browsing sessions for data collection and archival purposes. Notable features include its browser-based architecture and ability to run complex crawls efficiently, making it suitable for projects requiring comprehensive web data capture.