serp.fast

Open source

Notable repositories

Curated selection of search, extraction, and browser automation repositories. Filter by category or language, or search across editorial context.

Turn websites into LLM-ready markdown with a single API call.

155K+TypeScriptAI Scraping

AI framework that gives LLMs the ability to control a web browser.

98K+PythonAI Scraping

Open-source async crawler optimized for LLM data extraction.

68K+PythonAI Scraping

A Python library from Google for extracting structured information from unstructured text using LLMs, with precise source grounding and interactive visualization.

36K+PythonAI Scraping

LLM-orchestrated web scraping with automatic graph-based extraction.

27K+PythonAI Scraping

Natural language browser automation built on Playwright.

23K+TypeScriptAI Scraping

An open-source, self-hosted no-code platform for web scraping, crawling, and AI data extraction that turns websites into structured APIs.

15K+TypeScriptAI Scraping

A service that converts any URL into clean, LLM-friendly Markdown by prefixing it with r.jina.ai, and can be self-hosted.

11K+TypeScriptAI Scraping

A TypeScript library that turns any webpage into structured data using LLMs, with Zod schemas defining the output and Playwright handling the browser.

6.8K+TypeScriptAI Scraping

A Python framework for LLM-based structured extraction from documents, aimed at minimal boilerplate.

1.8K+PythonAI Scraping

Fast, local-first web content extraction for LLMs. CLI, REST API, and MCP server – built in Rust.

1.8K+RustAI Scraping

A lightweight Python library for scraping websites with LLMs using minimal code and token-efficient prompts, built on Playwright.

1.2K+PythonAI Scraping

A fully local autonomous agent that browses the web and writes code without API keys or cloud bills.

26K+PythonAI Web Agents

An LLM- and computer-vision-driven agent that automates browser workflows across websites without per-site selectors.

21K+PythonAI Web Agents

An open-source Chrome extension that runs multi-agent web-automation workflows using your own LLM API key.

13K+TypeScriptAI Web Agents

An open-source agentic browser, built on Chromium, positioned against ChatGPT Atlas, Perplexity Comet, and Dia.

11K+TypeScriptAI Web Agents

An experimental Microsoft Research agent that operates across the browser and the local file system.

9.9K+PythonAI Web Agents

A Large Action Model framework for building web agents that turn natural-language goals into Selenium or Playwright actions.

6.3K+PythonAI Web Agents

An open-source, vision-first browser agent for the TypeScript and Playwright ecosystem.

4.1K+TypeScriptAI Web Agents

An open-source framework for building web agents and deploying serverless web-automation functions on managed browser infrastructure.

1.9K+PythonAI Web Agents

Vision utilities from Reworkd, including element tagging and OCR, that make webpages legible to multimodal LLM agents.

1.7K+Jupyter NotebookAI Web Agents

An AI-native browser-automation framework from Hyperbrowser that extends Playwright with natural-language commands.

1.4K+TypeScriptAI Web Agents

A hierarchical browser-automation agent from Emergence AI, used as a reference web agent on the WebVoyager benchmark.

1.2K+PythonAI Web Agents

A fully private, open-source browser assistant that runs models on-device for in-page tasks.

1K+TypeScriptAI Web Agents

An open-source, AI-powered browser-assistant extension that drives the page through natural language.

978TypeScriptAI Web Agents

Proxy server that solves Cloudflare and DDoS-Guard challenges with a real headless browser and hands back the cookies and user-agent so plain HTTP clients can get through.

14K+PythonAnti-Detection

Custom anti-detect build of Firefox with Playwright bindings that spoofs browser fingerprints at the C++ level.

9.2K+C++Anti-Detection

Plugin framework for Puppeteer (and Playwright via playwright-extra) with stealth and ad-blocking plugins.

7.3K+TypeScriptAnti-Detection

Python requests wrapper that bypasses Cloudflare's JavaScript anti-bot interstitial pages.

6.6K+PythonAnti-Detection

curl with browser TLS fingerprints to bypass anti-bot detection.

6K+CAnti-Detection

Python HTTP client that binds curl-impersonate to mimic real browser TLS, JA3, and HTTP/2 fingerprints.

5.8K+PythonAnti-Detection

Python web automation without a traditional webdriver dependency.

4.4K+PythonAnti-Detection

Patched, drop-in replacement for Playwright that removes the CDP and runtime leaks bot detectors look for.

3.5K+TypeScriptAnti-Detection

utls

Fork of Go's standard TLS library that exposes low-level control over the ClientHello for fingerprint mimicry.

2.4K+GoAnti-Detection

Apify's TypeScript toolkit that generates and injects realistic, internally consistent browser fingerprints into Playwright and Puppeteer.

2.4K+TypeScriptAnti-Detection

Lightweight script that drives a real browser via DrissionPage to pass Cloudflare verification for scraping.

2.4K+PythonAnti-Detection

surf

Go HTTP client with Chrome and Firefox impersonation, HTTP/3 QUIC fingerprinting, and JA3/JA4 TLS emulation.

1.7K+GoAnti-Detection

Go HTTP client built on utls that spoofs browser TLS, JA3, and HTTP/2 fingerprints, with bindings for other languages.

1.7K+GoAnti-Detection

Puppeteer launcher that behaves like a real browser to clear Cloudflare and similar bot-detection captchas while keeping the standard Puppeteer API.

1.6K+JavaScriptAnti-Detection

Library that spoofs TLS and JA3 fingerprints from both Go and JavaScript.

1.4K+GoAnti-Detection

Patches for Puppeteer and Playwright that strip the CDP and runtime leak signals (such as the Runtime.Enable tell) used to fingerprint automation.

1.4K+JavaScriptAnti-Detection

Async-first, CDP-based undetectable web-automation framework forked from nodriver, with Docker support.

1.3K+PythonAnti-Detection

Chrome DevTools Protocol automation for Node.js.

94K+TypeScriptBrowser Automation

Cross-browser automation library for Chromium, Firefox, and WebKit.

90K+TypeScriptBrowser Automation

A browser automation CLI from Vercel Labs, written in Rust, for AI agents to drive a real browser.

36K+RustBrowser Automation

Browser automation framework supporting multiple languages and browsers.

34K+JavaBrowser Automation

A headless browser written from scratch in Zig for AI and automation workloads, speaking the Chrome DevTools Protocol.

31K+ZigBrowser Automation

The official Python bindings for Playwright, automating Chromium, Firefox and WebKit through one API.

14K+PythonBrowser Automation

Dockerized headless-browser infrastructure that exposes Puppeteer and Playwright over a web service.

13K+TypeScriptBrowser Automation

An idiomatic Go package for driving Chrome DevTools Protocol browsers with no external dependencies.

13K+GoBrowser Automation

Custom Selenium chromedriver that avoids detection by anti-bot services.

12K+PythonBrowser Automation

A Python framework for UI testing, web scraping and stealth automation built on Selenium with a CDP-based undetected mode.

12K+PythonBrowser Automation

An open-source browser API for AI agents with built-in session management, proxies and CAPTCHA handling.

7.2K+TypeScriptBrowser Automation

A Chrome DevTools Protocol driver for Go offering high-level web automation and scraping with auto-waiting.

7K+GoBrowser Automation

An async Python library that automates Chromium without a WebDriver, with native CAPTCHA bypass and realistic interactions.

6.9K+PythonBrowser Automation

A Node.js end-to-end testing framework with one high-level API over Playwright, Puppeteer and WebDriver backends.

4.2K+JavaScriptBrowser Automation

A lightweight, scriptable browser-as-a-service with an HTTP API for JavaScript rendering in scraping pipelines.

4.1K+PythonBrowser Automation

An unofficial Python port of Puppeteer for controlling headless Chromium over the DevTools Protocol.

3.9K+PythonBrowser Automation

The official .NET port of Puppeteer for driving headless Chromium and Chrome from C#.

3.9K+C#Browser Automation

A library that runs a pool of parallel Puppeteer instances with queuing, retries and error handling.

3.5K+TypeScriptBrowser Automation

A pretrained, training-free OCR and object-detection model for recognizing text and slider CAPTCHAs, packaged for pip.

14K+PythonCAPTCHA Solving

A browser extension that solves reCAPTCHA, hCaptcha, FunCaptcha, Turnstile and text CAPTCHAs, with hooks for Selenium, Puppeteer and Playwright.

10K+N/ACAPTCHA Solving

A Chrome, Edge and Firefox extension that solves reCAPTCHA by running its audio challenge through speech-to-text.

9.1K+JavaScriptCAPTCHA Solving

A TensorFlow framework using CNN/ResNet/DenseNet with GRU/LSTM and CTC to train custom image-CAPTCHA recognition models.

3.2K+PythonCAPTCHA Solving

A Python library that solves hCaptcha image challenges using multimodal LLMs and YOLO models, usable from Playwright.

2.3K+PythonCAPTCHA Solving

A Python library that solves reCAPTCHA v2 and v3 via the audio-challenge speech-to-text approach, with DrissionPage and Selenium support.

1.8K+PythonCAPTCHA Solving

A Python solver that obtains Cloudflare Turnstile tokens through Patchright/Playwright browser automation and exposes them via an API server.

829PythonCAPTCHA Solving

A self-hosted Ruby platform for building agents that monitor the web, scrape pages, watch feeds, and act or notify on changes.

49K+RubyChange Detection

An open-source feed generator that turns thousands of sites without RSS into feeds through per-site route adapters.

44K+TypeScriptChange Detection

A self-hosted tool for website change detection and monitoring, with text, XPath, and JSON diffing, restock and price-drop alerts, and notifications across many channels.

31K+PythonChange Detection

A self-hosted PHP service that generates RSS, Atom, and JSON feeds for sites that lack them, using maintained per-site bridges.

9K+PHPChange Detection

A configurable command-line tool that watches parts of webpages or command output and notifies you via email, Telegram, and other channels when something changes.

3.1K+PythonChange Detection

Crawls a site from a starting URL and bundles the content into a knowledge file for building a custom GPT.

22K+TypeScriptCrawlers & Search

A Go crawling and spidering framework with headless and JavaScript-aware modes.

17K+GoCrawlers & Search

Redis-based components that give Scrapy a shared request queue for distributed crawling across workers.

5.6K+PythonCrawlers & Search

A fast Go crawler for discovering endpoints, assets, and JavaScript sources in a web application.

5K+GoCrawlers & Search

YaCy

A decentralized peer-to-peer search engine with its own crawler and index, designed to run without a central server.

3.9K+JavaCrawlers & Search

A distributed crawler management framework built on Scrapy, Scrapyd, Django, and Vue.js, with a web dashboard for deploying and monitoring spiders.

3.5K+PythonCrawlers & Search

The Internet Archive's open-source, extensible web crawler built for web-scale, archival-quality capture.

3.2K+JavaCrawlers & Search

An extensible, scalable open-source web crawler in Java, built to run on Hadoop.

3.2K+JavaCrawlers & Search

A concurrent PHP crawler library built on Guzzle that can execute JavaScript via headless Chrome.

2.8K+PHPCrawlers & Search

A low-latency Rust web crawler and data collector with headless rendering and LLM-ready output.

2.5K+RustCrawlers & Search

An integrated Python crawler and extractor that pulls structured article text and metadata from news sites.

2.4K+PythonCrawlers & Search

pywb

A Python web archiving toolkit for recording and replaying WARC and WACZ web archives.

1.6K+JavaScriptCrawlers & Search

A configurable, extensible PHP web spider with depth- and breadth-first traversal, URL filtering, and pluggable discovery and persistence.

1.3K+PHPCrawlers & Search

A Python tool from Microsoft for converting files and Office documents (HTML, PDF, Word, Excel, PowerPoint) to Markdown for LLM ingestion.

153K+PythonData Parsing

A document parsing toolkit that converts PDF, HTML, and DOCX into structured Markdown or JSON for RAG and LLM pipelines.

61K+PythonData Parsing

Fast, flexible jQuery-like HTML parser for Node.js.

30K+TypeScriptData Parsing

newspaper3k, a Python 3 library for extracting full text, article metadata, and news content from web pages.

15K+PythonData Parsing

A Go library that brings jQuery-style HTML parsing and selection to Go.

14K+GoData Parsing

A Java HTML parser with DOM traversal, CSS-selector extraction, and HTML cleaning for XSS safety.

11K+JavaData Parsing

A standalone JavaScript version of the article-extraction algorithm behind Firefox's Reader View.

11K+JavaScriptData Parsing

A libxml2-backed Ruby library for parsing HTML and XML with XPath and CSS selectors.

6.2K+CData Parsing

A Python library and command-line tool for gathering text and metadata from web pages and feeds, with output to CSV, JSON, HTML, Markdown, TXT, and XML.

6.1K+PythonData Parsing

A .NET library that parses HTML5, MathML, SVG, and CSS into a W3C-spec DOM, queryable with LINQ and CSS selectors.

5.5K+C#Data Parsing

A fast, forgiving streaming HTML and XML parser for Node.js.

4.7K+TypeScriptData Parsing

A Go command-line tool for scraping and extracting data from web pages and JSON using HTML, CSS, and JSON selectors.

4.7K+GoData Parsing

A Python library that extracts the main article body, title, and lead image from HTML pages.

4.1K+HTMLData Parsing

A WHATWG HTML5 spec-compliant HTML parsing and serialization toolset for Node.js.

3.9K+TypeScriptData Parsing

A Go library and CLI that converts HTML into clean Markdown, with rule-based extensibility and support for entire websites.

3.7K+GoData Parsing

lxml

High-performance XML and HTML processing library for Python.

3K+PythonData Parsing

Python library for pulling data out of HTML and XML files.

N/APythonData Parsing

The official Model Context Protocol monorepo of reference servers, including the canonical Fetch server (URL to Markdown) and a Puppeteer browser-automation server.

87K+TypeScriptMCP Servers

The Chrome team's MCP server that exposes Chrome DevTools to coding agents for browsing, performance tracing, and debugging live web apps.

43K+TypeScriptMCP Servers

Microsoft's official MCP server that lets agents drive a browser through structured accessibility-tree snapshots instead of screenshots.

33K+TypeScriptMCP Servers

A Chrome extension-based MCP server that exposes your real logged-in browser to AI assistants for automation, content analysis, and semantic search.

11K+TypeScriptMCP Servers

An MCP server that connects AI applications to your existing local browser through an extension, reusing real sessions and cookies.

6.7K+TypeScriptMCP Servers

The official Firecrawl MCP server that adds web scraping, crawling, and search tools to Cursor, Claude, and other MCP clients.

6.6K+JavaScriptMCP Servers

An MCP server that converts web pages, PDFs, images, and documents into clean Markdown for LLM ingestion.

2.7K+TypeScriptMCP Servers

Tavily's official MCP server providing agents with real-time search, extract, map, and crawl tools tuned for LLM consumption.

2.1K+JavaScriptMCP Servers

A multi-engine MCP server, CLI, and local daemon that runs agent web search across engines like DuckDuckGo, Bing, and Brave with no API keys.

1.6K+TypeScriptMCP Servers

Apify's MCP server that exposes thousands of Apify Actors (ready-made scrapers and crawlers) as callable tools for AI agents.

2.1K+TypeScriptMCP Servers

A lightweight MCP server providing DuckDuckGo web search plus page-content fetching, with no API key required.

1.2K+PythonMCP Servers

An MCP server that wraps the browser-use agent in Docker (with a VNC view) so any MCP client can run autonomous browser tasks.

823PythonMCP Servers

A flexible HTTP fetching MCP server that returns web content as HTML, JSON, plain text, or Markdown with custom headers.

781TypeScriptMCP Servers

An interactive, TLS-capable intercepting HTTP/HTTPS proxy with a scriptable Python API.

43K+PythonProxy & Networking

An async Python finder, checker, and server for free public HTTP(S) and SOCKS proxies, including a rotating proxy server mode.

4.1K+PythonProxy & Networking

A self-hosted proxy pool that scrapes, validates, and serves free proxies through a local HTTP API.

4K+PythonProxy & Networking

A lightweight, zero-dependency, pluggable HTTP/HTTPS proxy server framework in Python with TLS interception.

3.5K+PythonProxy & Networking

A Go proxy checker and IP rotator that can run as a rotating proxy server in front of a scraper.

2.1K+GoProxy & Networking

An async Rust tool that scrapes and checks HTTP, SOCKS4, and SOCKS5 proxies with filtering and flexible output.

1.2K+RustProxy & Networking

A Node.js monorepo of HTTP, HTTPS, and SOCKS proxy agents, including https-proxy-agent.

1.1K+TypeScriptProxy & Networking

A Node.js proxy server with SSL, HTTP/HTTPS, SOCKS5, authentication, and upstream proxy chaining.

1K+JavaScriptProxy & Networking

A Scrapy downloader middleware that rotates requests across a list of proxies and bans dead ones.

773PythonProxy & Networking

Adaptive web scraping framework with smart element tracking, anti-bot bypass, and stealth browser mode.

63K+PythonScraping Frameworks

Fast, high-level web crawling and scraping framework for Python.

62K+PythonScraping Frameworks

Elegant scraping framework for Go with a clean callback API.

25K+GoScraping Frameworks

Web scraping and browser automation library for Node.js.

23K+TypeScriptScraping Frameworks

A distributed Python web crawler system with a web UI, scheduler, script editor and result viewer.

16K+PythonScraping Frameworks

A Pythonic HTML parsing layer over requests with CSS and XPath selectors plus JavaScript rendering via pyppeteer.

13K+PythonScraping Frameworks

A scalable, modular web crawler framework for Java built around a downloader, scheduler and pipeline architecture.

11K+JavaScraping Frameworks

Apify's Python crawling framework that unifies HTTP and headless-browser scraping with auto-scaling, proxy rotation and request queues.

9.2K+PythonScraping Frameworks

A lightweight Python scraper that learns extraction rules from example data you provide.

7.2K+PythonScraping Frameworks

A Node.js web crawler with server-side jQuery via Cheerio, plus built-in rate limiting, retries and request queueing.

6.7K+TypeScriptScraping Frameworks

A Go-based declarative data extraction engine with its own FQL query language covering both static and browser-rendered pages.

6K+GoScraping Frameworks

A declarative Node.js scraper with a composable selector DSL that follows pagination and streams results to files or databases.

5.9K+JavaScriptScraping Frameworks

Python library for automating interaction with websites.

4.8K+PythonScraping Frameworks

An all-in-one Python scraping framework with built-in anti-detection features aimed at bypassing Cloudflare and similar bot mitigation.

5.5K+PythonScraping Frameworks

A Ruby library that automates stateful website interaction including forms, links, cookies and history.

4.4K+RubyScraping Frameworks

Scrapy integration for the Splash JavaScript-rendering headless browser service.

3.2K+PythonScraping Frameworks

A concurrent Go web crawling and scraping framework with JS rendering, caching and Scrapy-like middleware pipelines.

2.7K+GoScraping Frameworks

A Node.js tool that downloads an entire website to a local directory, including CSS, images and JS so pages render offline.

1.7K+JavaScriptScraping Frameworks

A complete Scrapy-inspired web scraping toolkit for PHP with spiders, middleware and item pipelines.

1.4K+PHPScraping Frameworks