The best open-source Data Scraping tools, curated: 25 in the registry, including Browser Harness, browser-use, browser-use web-ui. Every tool is open-source; each entry lists license, self-hostability, and links.
Browser Harness— tool, 2026-09-03 Pipeline inventory: Stage 5 strong acquisition fallback (fit 8/10) and Stage 2 partial enrichment extractor (fit 6/10). Self-healing harness connects LLMs to a real Chrome browser via CDP/Playwright and writes reusable helpers at runtime; complements Agent Reach/Crawl4AI/Crawlee/Lightpanda; ~16.7k stars. Expected to reduce interactive public-web extraction gaps by roughly 40-60%, but does not supply proprietary firmographic/contact data. CAVEATS: target-site ToS and…
browser-use— tool, 2026-09-03 Flagship 'make websites accessible for AI agents'; ~108k stars; CDP/Playwright; complements Crawl4AI/Browser Harness/Lightpanda; CAVEAT target-site ToS + agent runtime actions
browser-use web-ui— tool, 2026-09-03 Web GUI to run browser-use agents; ~16k stars; complements browser-use
cdp-use— tool, 2026-09-03 Pure Chrome DevTools Protocol, type-safe in Python; browser-automation primitive; ~308 stars
Crawl4AI— tool, 2026-09-03 OSS alt to Browserbase; ~77347 stars; via openalternative.co
Darts— tool, 2026-09-03 ~9487 stars; via ossinsight.io
dlt— tool, 2026-09-03 Pure OSI Python EL library; schema inference
OpenSERP— tool, 2026-09-03 OSS alt to SearchAPI; ~1235 stars; via openalternative.co
PyOD— tool, 2026-09-03 ~9953 stars; via ossinsight.io
ScrapeGraphAI— tool, 2026-09-03 OSS alt to Diffbot; ~29238 stars; via opensource.builders
Scrapy— tool, 2026-09-03 Mature; Zyte-backed; no built-in JS render
spaCy— tool, 2026-09-03 Production NLP; Explosion-maintained
Stanza— tool, 2026-09-03 SOTA neural accuracy; 60+ languages
STUMPY— tool, 2026-09-03 ~4140 stars; via ossinsight.io
Firecrawl anydoc— tool, 2026-09-03 Open-source document-to-Markdown converter from the Firecrawl team (Rust with multi-language bindings): Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF into clean GitHub-Flavoured Markdown — an ingest layer for agent and RAG pipelines.
Obscura— tool, 2026-08-31 Open-source headless browser for AI agents and web scraping, built in Rust: runs V8 JavaScript, speaks Chrome DevTools Protocol, drop-in for Puppeteer/Playwright, ~30MB, with stealth/anti-fingerprinting and MCP support — no Chromium.
Maxun— tool, 2026-08-31 Open-source, no-code web-data platform: extract structured data (recorded actions or natural language), scrape pages to Markdown/HTML with screenshots, crawl whole sites, and run automated searches — self-hostable, with an SDK and CLI.
Chrome DevTools MCP— tool, 2026-08-25 MCP server that lets an agent inspect and control a live Chrome browser — performance traces, network requests, console errors, rendered pages, and screenshots.
Docling— tool, 2026-08-25 Document-processing toolkit with advanced PDF/DOCX/PPTX/XLSX understanding (tables, complex layouts) and generative-AI integrations — for when simple PDF-to-text loses structure.
Firecrawl— tool, 2026-08-25 The context API to search, scrape, and crawl the web at scale and turn sites into clean, LLM-ready data. Core is AGPL-3.0; SDKs/UI are MIT.
Firecrawl MCP Server— tool, 2026-08-25 Official Firecrawl MCP server that brings web search, scraping, crawling, and structured extraction to MCP-compatible agents as clean, agent-ready context.
MarkItDown— tool, 2026-08-25 Microsoft utility that converts PDFs, Word, PowerPoint, Excel, HTML, images and more into clean Markdown for LLM pipelines — structure-preserving.
Playwright MCP— tool, 2026-08-25 Microsoft MCP server that lets agents automate a browser via structured accessibility snapshots (not screenshots) — for rendered-page checks, forms, funnels, and landing-page QA.
oc (only-cli)— tool, 2026-08-25 Open-source CLI that fetches a web page and returns a compact numbered view instead of raw HTML — a cheap web-read layer for scraping/enrichment agents. Installs as a CLI and as an agent skill (npx skills add).