Extracts structured content from web pages. Respects rate limits and robots.txt.
Sign in to vote
A scraping skill with built-in politeness: it reads and respects `robots.txt`, applies configurable rate limiting, rotates user-agent strings, and handles JavaScript-rendered pages via a bundled Playwright headless browser mode. Returns structured data (title, main content, metadata, links) rather than raw HTML.
In static mode, uses a lightweight HTTP client with HTML parsing. In JavaScript mode, spins up a Playwright browser instance with a pre-configured profile that minimizes detection. Robots.txt is fetched once per domain and cached for the session lifetime. Rate limiting is implemented with a token-bucket algorithm per domain.