Reader API that turns any URL into clean markdown, now part of Elastic
CategoryAI-Native Search APIs Self-hosted metasearch that aggregates results from 70+ engines
CategoryAI-Native Search APIs Turns any website into LLM-ready markdown via a single API
CategoryWeb Crawl & Data Extraction APIs Adaptive Python scraper whose selectors survive site redesigns
CategoryWeb Crawl & Data Extraction APIs Rust page-to-markdown API that skips the headless browser
CategoryWeb Crawl & Data Extraction APIs Open-source browser API with stealth and session management
CategoryBrowser Infrastructure Zig-built headless browser using far less memory than Chrome
CategoryBrowser Infrastructure Hands any website to an AI agent, the top browser-automation repo
CategoryBrowser Infrastructure Nonprofit web archive of 9.5 petabytes behind most major LLMs
CategoryIndependent Web Indexes Open-source search you re-rank yourself, built on its own index
CategoryIndependent Web Indexes Search engine for the text-heavy old web Google buries
CategoryIndependent Web Indexes Open, decentralized index aiming to rival Google and Bing
CategoryIndependent Web Indexes Long-running open-source search engine with its own web index
CategoryIndependent Web Indexes Python library that scrapes sites from plain-English prompts
CategoryAgentic Extraction TypeScript SDK for driving browsers with natural-language steps
CategoryAgentic Extraction AI agent that navigates and extracts using vision and LLMs
CategoryAgentic Extraction The original Python framework for large-scale web crawling
CategoryOpen Source Frameworks Microsoft's cross-browser automation for testing and scraping
CategoryOpen Source Frameworks Google's Node.js library for driving headless Chrome
CategoryOpen Source Frameworks Apify's crawling library wrapping Playwright and Puppeteer
CategoryOpen Source Frameworks Veteran browser automation with bindings for most languages
CategoryOpen Source Frameworks Open-source crawler shaped for RAG, the most-starred on GitHub
CategoryOpen Source Frameworks Python parser that turns messy HTML into navigable trees
CategoryOpen Source Frameworks jQuery-style HTML parsing for Node.js, no browser required
CategoryOpen Source Frameworks Fast concurrent scraping framework for Go with a callback API
CategoryOpen Source Frameworks Pairs Requests and Beautiful Soup for form-driven scraping
CategoryOpen Source Frameworks Lightweight async scraping from HTTPx plus Scrapy Parsel
CategoryOpen Source Frameworks Strips nav, ads, and chrome to leave the main article text
CategoryOpen Source Frameworks The library behind Firefox Reader Mode, extracting article text
CategoryOpen Source Frameworks Benchmark testing browser agents on 153 tasks across 144 live sites
812 long-horizon web tasks on reproducible self-hosted sites
643 live-web tasks measuring end-to-end multimodal agents
2,350 tasks across 137 sites measuring cross-domain transfer
369 computer-use tasks spanning Ubuntu, Windows, and macOS