Turn websites into LLM-ready markdown with a single API call.
155K+· TypeScript · AI Scraping AI framework that gives LLMs the ability to control a web browser.
Open-source async crawler optimized for LLM data extraction.
A Python library from Google for extracting structured information from unstructured text using LLMs, with precise source grounding and interactive visualization.
LLM-orchestrated web scraping with automatic graph-based extraction.
Natural language browser automation built on Playwright.
23K+· TypeScript · AI Scraping An open-source, self-hosted no-code platform for web scraping, crawling, and AI data extraction that turns websites into structured APIs.
15K+· TypeScript · AI Scraping A service that converts any URL into clean, LLM-friendly Markdown by prefixing it with r.jina.ai, and can be self-hosted.
11K+· TypeScript · AI Scraping A TypeScript library that turns any webpage into structured data using LLMs, with Zod schemas defining the output and Playwright handling the browser.
6.8K+· TypeScript · AI Scraping A Python framework for LLM-based structured extraction from documents, aimed at minimal boilerplate.
Fast, local-first web content extraction for LLMs. CLI, REST API, and MCP server – built in Rust.
A lightweight Python library for scraping websites with LLMs using minimal code and token-efficient prompts, built on Playwright.
A fully local autonomous agent that browses the web and writes code without API keys or cloud bills.
26K+· Python · AI Web Agents An LLM- and computer-vision-driven agent that automates browser workflows across websites without per-site selectors.
21K+· Python · AI Web Agents An open-source Chrome extension that runs multi-agent web-automation workflows using your own LLM API key.
13K+· TypeScript · AI Web Agents An open-source agentic browser, built on Chromium, positioned against ChatGPT Atlas, Perplexity Comet, and Dia.
11K+· TypeScript · AI Web Agents An experimental Microsoft Research agent that operates across the browser and the local file system.
9.9K+· Python · AI Web Agents A Large Action Model framework for building web agents that turn natural-language goals into Selenium or Playwright actions.
6.3K+· Python · AI Web Agents An open-source, vision-first browser agent for the TypeScript and Playwright ecosystem.
4.1K+· TypeScript · AI Web Agents An open-source framework for building web agents and deploying serverless web-automation functions on managed browser infrastructure.
1.9K+· Python · AI Web Agents Vision utilities from Reworkd, including element tagging and OCR, that make webpages legible to multimodal LLM agents.
1.7K+· Jupyter Notebook · AI Web Agents An AI-native browser-automation framework from Hyperbrowser that extends Playwright with natural-language commands.
1.4K+· TypeScript · AI Web Agents A hierarchical browser-automation agent from Emergence AI, used as a reference web agent on the WebVoyager benchmark.
1.2K+· Python · AI Web Agents A fully private, open-source browser assistant that runs models on-device for in-page tasks.
1K+· TypeScript · AI Web Agents An open-source, AI-powered browser-assistant extension that drives the page through natural language.
978· TypeScript · AI Web Agents Proxy server that solves Cloudflare and DDoS-Guard challenges with a real headless browser and hands back the cookies and user-agent so plain HTTP clients can get through.
14K+· Python · Anti-Detection Custom anti-detect build of Firefox with Playwright bindings that spoofs browser fingerprints at the C++ level.
Plugin framework for Puppeteer (and Playwright via playwright-extra) with stealth and ad-blocking plugins.
7.3K+· TypeScript · Anti-Detection Python requests wrapper that bypasses Cloudflare's JavaScript anti-bot interstitial pages.
6.6K+· Python · Anti-Detection curl with browser TLS fingerprints to bypass anti-bot detection.
Python HTTP client that binds curl-impersonate to mimic real browser TLS, JA3, and HTTP/2 fingerprints.
5.8K+· Python · Anti-Detection Python web automation without a traditional webdriver dependency.
4.4K+· Python · Anti-Detection Patched, drop-in replacement for Playwright that removes the CDP and runtime leaks bot detectors look for.
3.5K+· TypeScript · Anti-Detection Fork of Go's standard TLS library that exposes low-level control over the ClientHello for fingerprint mimicry.
Apify's TypeScript toolkit that generates and injects realistic, internally consistent browser fingerprints into Playwright and Puppeteer.
2.4K+· TypeScript · Anti-Detection Lightweight script that drives a real browser via DrissionPage to pass Cloudflare verification for scraping.
2.4K+· Python · Anti-Detection Go HTTP client with Chrome and Firefox impersonation, HTTP/3 QUIC fingerprinting, and JA3/JA4 TLS emulation.
Go HTTP client built on utls that spoofs browser TLS, JA3, and HTTP/2 fingerprints, with bindings for other languages.
Puppeteer launcher that behaves like a real browser to clear Cloudflare and similar bot-detection captchas while keeping the standard Puppeteer API.
1.6K+· JavaScript · Anti-Detection Library that spoofs TLS and JA3 fingerprints from both Go and JavaScript.
Patches for Puppeteer and Playwright that strip the CDP and runtime leak signals (such as the Runtime.Enable tell) used to fingerprint automation.
1.4K+· JavaScript · Anti-Detection Async-first, CDP-based undetectable web-automation framework forked from nodriver, with Docker support.
1.3K+· Python · Anti-Detection Chrome DevTools Protocol automation for Node.js.
94K+· TypeScript · Browser Automation Cross-browser automation library for Chromium, Firefox, and WebKit.
90K+· TypeScript · Browser Automation A browser automation CLI from Vercel Labs, written in Rust, for AI agents to drive a real browser.
36K+· Rust · Browser Automation Browser automation framework supporting multiple languages and browsers.
34K+· Java · Browser Automation A headless browser written from scratch in Zig for AI and automation workloads, speaking the Chrome DevTools Protocol.
31K+· Zig · Browser Automation The official Python bindings for Playwright, automating Chromium, Firefox and WebKit through one API.
14K+· Python · Browser Automation Dockerized headless-browser infrastructure that exposes Puppeteer and Playwright over a web service.
13K+· TypeScript · Browser Automation An idiomatic Go package for driving Chrome DevTools Protocol browsers with no external dependencies.
13K+· Go · Browser Automation Custom Selenium chromedriver that avoids detection by anti-bot services.
12K+· Python · Browser Automation A Python framework for UI testing, web scraping and stealth automation built on Selenium with a CDP-based undetected mode.
12K+· Python · Browser Automation An open-source browser API for AI agents with built-in session management, proxies and CAPTCHA handling.
7.2K+· TypeScript · Browser Automation A Chrome DevTools Protocol driver for Go offering high-level web automation and scraping with auto-waiting.
7K+· Go · Browser Automation An async Python library that automates Chromium without a WebDriver, with native CAPTCHA bypass and realistic interactions.
6.9K+· Python · Browser Automation A Node.js end-to-end testing framework with one high-level API over Playwright, Puppeteer and WebDriver backends.
4.2K+· JavaScript · Browser Automation A lightweight, scriptable browser-as-a-service with an HTTP API for JavaScript rendering in scraping pipelines.
4.1K+· Python · Browser Automation An unofficial Python port of Puppeteer for controlling headless Chromium over the DevTools Protocol.
3.9K+· Python · Browser Automation The official .NET port of Puppeteer for driving headless Chromium and Chrome from C#.
3.9K+· C# · Browser Automation A library that runs a pool of parallel Puppeteer instances with queuing, retries and error handling.
3.5K+· TypeScript · Browser Automation A pretrained, training-free OCR and object-detection model for recognizing text and slider CAPTCHAs, packaged for pip.
14K+· Python · CAPTCHA Solving A browser extension that solves reCAPTCHA, hCaptcha, FunCaptcha, Turnstile and text CAPTCHAs, with hooks for Selenium, Puppeteer and Playwright.
A Chrome, Edge and Firefox extension that solves reCAPTCHA by running its audio challenge through speech-to-text.
9.1K+· JavaScript · CAPTCHA Solving A TensorFlow framework using CNN/ResNet/DenseNet with GRU/LSTM and CTC to train custom image-CAPTCHA recognition models.
3.2K+· Python · CAPTCHA Solving A Python library that solves hCaptcha image challenges using multimodal LLMs and YOLO models, usable from Playwright.
2.3K+· Python · CAPTCHA Solving A Python library that solves reCAPTCHA v2 and v3 via the audio-challenge speech-to-text approach, with DrissionPage and Selenium support.
1.8K+· Python · CAPTCHA Solving A Python solver that obtains Cloudflare Turnstile tokens through Patchright/Playwright browser automation and exposes them via an API server.
829· Python · CAPTCHA Solving A self-hosted Ruby platform for building agents that monitor the web, scrape pages, watch feeds, and act or notify on changes.
49K+· Ruby · Change Detection An open-source feed generator that turns thousands of sites without RSS into feeds through per-site route adapters.
44K+· TypeScript · Change Detection A self-hosted tool for website change detection and monitoring, with text, XPath, and JSON diffing, restock and price-drop alerts, and notifications across many channels.
31K+· Python · Change Detection A self-hosted PHP service that generates RSS, Atom, and JSON feeds for sites that lack them, using maintained per-site bridges.
A configurable command-line tool that watches parts of webpages or command output and notifies you via email, Telegram, and other channels when something changes.
3.1K+· Python · Change Detection Crawls a site from a starting URL and bundles the content into a knowledge file for building a custom GPT.
22K+· TypeScript · Crawlers & Search A Go crawling and spidering framework with headless and JavaScript-aware modes.
17K+· Go · Crawlers & Search Redis-based components that give Scrapy a shared request queue for distributed crawling across workers.
5.6K+· Python · Crawlers & Search A fast Go crawler for discovering endpoints, assets, and JavaScript sources in a web application.
A decentralized peer-to-peer search engine with its own crawler and index, designed to run without a central server.
3.9K+· Java · Crawlers & Search A distributed crawler management framework built on Scrapy, Scrapyd, Django, and Vue.js, with a web dashboard for deploying and monitoring spiders.
3.5K+· Python · Crawlers & Search The Internet Archive's open-source, extensible web crawler built for web-scale, archival-quality capture.
3.2K+· Java · Crawlers & Search An extensible, scalable open-source web crawler in Java, built to run on Hadoop.
3.2K+· Java · Crawlers & Search A concurrent PHP crawler library built on Guzzle that can execute JavaScript via headless Chrome.
2.8K+· PHP · Crawlers & Search A low-latency Rust web crawler and data collector with headless rendering and LLM-ready output.
2.5K+· Rust · Crawlers & Search An integrated Python crawler and extractor that pulls structured article text and metadata from news sites.
2.4K+· Python · Crawlers & Search A Python web archiving toolkit for recording and replaying WARC and WACZ web archives.
1.6K+· JavaScript · Crawlers & Search A configurable, extensible PHP web spider with depth- and breadth-first traversal, URL filtering, and pluggable discovery and persistence.
1.3K+· PHP · Crawlers & Search A Python tool from Microsoft for converting files and Office documents (HTML, PDF, Word, Excel, PowerPoint) to Markdown for LLM ingestion.
153K+· Python · Data Parsing A document parsing toolkit that converts PDF, HTML, and DOCX into structured Markdown or JSON for RAG and LLM pipelines.
Fast, flexible jQuery-like HTML parser for Node.js.
30K+· TypeScript · Data Parsing newspaper3k, a Python 3 library for extracting full text, article metadata, and news content from web pages.
A Go library that brings jQuery-style HTML parsing and selection to Go.
A Java HTML parser with DOM traversal, CSS-selector extraction, and HTML cleaning for XSS safety.
A standalone JavaScript version of the article-extraction algorithm behind Firefox's Reader View.
11K+· JavaScript · Data Parsing A libxml2-backed Ruby library for parsing HTML and XML with XPath and CSS selectors.
A Python library and command-line tool for gathering text and metadata from web pages and feeds, with output to CSV, JSON, HTML, Markdown, TXT, and XML.
6.1K+· Python · Data Parsing A .NET library that parses HTML5, MathML, SVG, and CSS into a W3C-spec DOM, queryable with LINQ and CSS selectors.
A fast, forgiving streaming HTML and XML parser for Node.js.
4.7K+· TypeScript · Data Parsing A Go command-line tool for scraping and extracting data from web pages and JSON using HTML, CSS, and JSON selectors.
A Python library that extracts the main article body, title, and lead image from HTML pages.
A WHATWG HTML5 spec-compliant HTML parsing and serialization toolset for Node.js.
3.9K+· TypeScript · Data Parsing A Go library and CLI that converts HTML into clean Markdown, with rule-based extensibility and support for entire websites.
High-performance XML and HTML processing library for Python.
Python library for pulling data out of HTML and XML files.
The official Model Context Protocol monorepo of reference servers, including the canonical Fetch server (URL to Markdown) and a Puppeteer browser-automation server.
87K+· TypeScript · MCP Servers The Chrome team's MCP server that exposes Chrome DevTools to coding agents for browsing, performance tracing, and debugging live web apps.
43K+· TypeScript · MCP Servers Microsoft's official MCP server that lets agents drive a browser through structured accessibility-tree snapshots instead of screenshots.
33K+· TypeScript · MCP Servers A Chrome extension-based MCP server that exposes your real logged-in browser to AI assistants for automation, content analysis, and semantic search.
11K+· TypeScript · MCP Servers An MCP server that connects AI applications to your existing local browser through an extension, reusing real sessions and cookies.
6.7K+· TypeScript · MCP Servers The official Firecrawl MCP server that adds web scraping, crawling, and search tools to Cursor, Claude, and other MCP clients.
6.6K+· JavaScript · MCP Servers An MCP server that converts web pages, PDFs, images, and documents into clean Markdown for LLM ingestion.
2.7K+· TypeScript · MCP Servers Tavily's official MCP server providing agents with real-time search, extract, map, and crawl tools tuned for LLM consumption.
2.1K+· JavaScript · MCP Servers A multi-engine MCP server, CLI, and local daemon that runs agent web search across engines like DuckDuckGo, Bing, and Brave with no API keys.
1.6K+· TypeScript · MCP Servers Apify's MCP server that exposes thousands of Apify Actors (ready-made scrapers and crawlers) as callable tools for AI agents.
2.1K+· TypeScript · MCP Servers A lightweight MCP server providing DuckDuckGo web search plus page-content fetching, with no API key required.
An MCP server that wraps the browser-use agent in Docker (with a VNC view) so any MCP client can run autonomous browser tasks.
A flexible HTTP fetching MCP server that returns web content as HTML, JSON, plain text, or Markdown with custom headers.
781· TypeScript · MCP Servers An interactive, TLS-capable intercepting HTTP/HTTPS proxy with a scriptable Python API.
43K+· Python · Proxy & Networking An async Python finder, checker, and server for free public HTTP(S) and SOCKS proxies, including a rotating proxy server mode.
4.1K+· Python · Proxy & Networking A self-hosted proxy pool that scrapes, validates, and serves free proxies through a local HTTP API.
4K+· Python · Proxy & Networking A lightweight, zero-dependency, pluggable HTTP/HTTPS proxy server framework in Python with TLS interception.
3.5K+· Python · Proxy & Networking A Go proxy checker and IP rotator that can run as a rotating proxy server in front of a scraper.
2.1K+· Go · Proxy & Networking An async Rust tool that scrapes and checks HTTP, SOCKS4, and SOCKS5 proxies with filtering and flexible output.
1.2K+· Rust · Proxy & Networking A Node.js monorepo of HTTP, HTTPS, and SOCKS proxy agents, including https-proxy-agent.
1.1K+· TypeScript · Proxy & Networking A Node.js proxy server with SSL, HTTP/HTTPS, SOCKS5, authentication, and upstream proxy chaining.
1K+· JavaScript · Proxy & Networking A Scrapy downloader middleware that rotates requests across a list of proxies and bans dead ones.
773· Python · Proxy & Networking Adaptive web scraping framework with smart element tracking, anti-bot bypass, and stealth browser mode.
63K+· Python · Scraping Frameworks Fast, high-level web crawling and scraping framework for Python.
62K+· Python · Scraping Frameworks Elegant scraping framework for Go with a clean callback API.
25K+· Go · Scraping Frameworks Web scraping and browser automation library for Node.js.
23K+· TypeScript · Scraping Frameworks A distributed Python web crawler system with a web UI, scheduler, script editor and result viewer.
16K+· Python · Scraping Frameworks A Pythonic HTML parsing layer over requests with CSS and XPath selectors plus JavaScript rendering via pyppeteer.
13K+· Python · Scraping Frameworks A scalable, modular web crawler framework for Java built around a downloader, scheduler and pipeline architecture.
11K+· Java · Scraping Frameworks Apify's Python crawling framework that unifies HTTP and headless-browser scraping with auto-scaling, proxy rotation and request queues.
9.2K+· Python · Scraping Frameworks A lightweight Python scraper that learns extraction rules from example data you provide.
7.2K+· Python · Scraping Frameworks A Node.js web crawler with server-side jQuery via Cheerio, plus built-in rate limiting, retries and request queueing.
6.7K+· TypeScript · Scraping Frameworks A Go-based declarative data extraction engine with its own FQL query language covering both static and browser-rendered pages.
6K+· Go · Scraping Frameworks A declarative Node.js scraper with a composable selector DSL that follows pagination and streams results to files or databases.
5.9K+· JavaScript · Scraping Frameworks Python library for automating interaction with websites.
4.8K+· Python · Scraping Frameworks An all-in-one Python scraping framework with built-in anti-detection features aimed at bypassing Cloudflare and similar bot mitigation.
5.5K+· Python · Scraping Frameworks A Ruby library that automates stateful website interaction including forms, links, cookies and history.
4.4K+· Ruby · Scraping Frameworks Scrapy integration for the Splash JavaScript-rendering headless browser service.
3.2K+· Python · Scraping Frameworks A concurrent Go web crawling and scraping framework with JS rendering, caching and Scrapy-like middleware pipelines.
2.7K+· Go · Scraping Frameworks A Node.js tool that downloads an entire website to a local directory, including CSS, images and JS so pages render offline.
1.7K+· JavaScript · Scraping Frameworks A complete Scrapy-inspired web scraping toolkit for PHP with spiders, middleware and item pipelines.
1.4K+· PHP · Scraping Frameworks