Cockroach Crawler
Built to be the best AI crawler for governed agents.
Cockroach Crawler is an open-source Node.js toolkit designed for web data acquisition, providing agents with controlled access to the internet. Key capabilities include:
• Static and rendered website crawling
• Adaptive traversal and URL mapping
• Page screenshots and PDF generation
• Structured data extraction (Markdown, JSON, JSONL)
• Explicit resource and runtime controls
This toolkit allows for comprehensive web data collection, ensuring that content from public GitHub, YouTube, and other web sources can be processed into evidence records. It supports various environments including MCP, Docker, and Cloudflare Workers, offering flexibility in deployment. The system prioritizes security and control, enforcing limits on origins, redirects, robots, requests, bytes, depth, and time, preventing unrestricted network access.
Cockroach Crawler is ideal for developers and organizations building specialized agents that require high-quality, governed web data. It provides the tools to crawl sites, map URLs, render JavaScript, and extract structured fields, transforming permitted public sources into ready-to-use formats while adhering to creator-defined policy limits. This ensures data integrity and responsible resource use, making it an excellent choice for research, content analysis, and data-driven applications.