{
  "$schema_doc": "https://gtmstacker.com/registry/schema/entry.schema.json",
  "stability": "emerging",
  "generator": "agentic-media-registry",
  "generated_at": "2026-09-26T00:00:00Z",
  "id": "com.gtmstacker.registry/tool/proxy-scraper",
  "type": "tool",
  "slug": "proxy-scraper",
  "canonical_url": "https://gtmstacker.com/registry/tool/proxy-scraper/",
  "title": "proxy-scraper",
  "description": "proxy-scraper is a self-hostable proxy scraper and verifier that pulls candidate proxies from 700+ sources, then runs honeypot, injected-script and TLS checks, and can serve a rotating proxy plus an MCP server. Open source: yes (MIT); self-hostable via pipx, pip or Docker. 5 stars, created 2026-09-24.",
  "category": "data-scraping",
  "tags": [
    "data-scraping",
    "mcp-agents",
    "self-hostable",
    "proxies",
    "verification",
    "rotating-proxy"
  ],
  "status": "active",
  "revision": 1,
  "content_hash": "df900df90b08156b8627866dfbe1f58a607b1d3aa3c14e843d38a1672eb9e981",
  "date_published": "2026-09-26T00:00:00Z",
  "date_modified": "2026-09-26T00:00:00Z",
  "source": {
    "name": "maximilianfeix · GitHub",
    "url": "https://github.com/maximilianfeix/proxy-scraper"
  },
  "license": "MIT",
  "one_liner": "proxy-scraper pulls proxies from 700+ sources and verifies them with honeypot, injected-script and TLS checks, then serves a rotating proxy and an MCP server.",
  "open_source": "yes",
  "self_hostable": "yes",
  "pricing_model": "free",
  "who_its_for": "An engineer running scraping or enrichment pipelines who wants owned proxy infrastructure — collection, verification and rotation they control — plus an MCP endpoint so an AI agent can fetch through vetted proxies, rather than paying a managed proxy vendor.",
  "aliases": [
    "proxy-scraper",
    "maximilianfeix/proxy-scraper",
    "proxy-scraper-mcp"
  ],
  "alternatives": [
    "deepscrape",
    "scrapy"
  ],
  "secondary_categories": [
    "mcp-agents"
  ],
  "last_verified": "2026-09-26",
  "evidence": {
    "claim_type": "mixed",
    "source_id": "https://github.com/maximilianfeix/proxy-scraper",
    "note": "MIT (OSI-open); 5 stars, created 2026-09-24, last push 2026-09-26. A self-hostable proxy scraper/verifier pulling from 700+ sources with honeypot, injected-script and TLS checks; ships a rotating proxy server and an MCP server for agents (WebFetch 2026-09-26, github.com/maximilianfeix/proxy-scraper). The README's 'one in five proxies injected a script' figure is the tool's own unverified claim, not backed by any dataset in the repo."
  },
  "caveats": "Facts verified from the repository README (WebFetch 2026-09-26); no independent testing. Vendor-claim, load-bearing: the README asserts 'one in five working proxies injected a script into a plain HTML page' but the repo ships no dataset or reproducible measurement behind it — treat that ratio as the tool's own unverified claim, not a fact. Free public proxies carry inherent trust and legal risk regardless of verification.",
  "lead": "proxy-scraper is a self-hostable proxy scraper and verifier that pulls candidate proxies from 700+ sources, then runs honeypot, injected-script and TLS checks, and can serve a rotating proxy plus an MCP server for AI agents. Open source: yes (MIT); self-hostable via pipx, pip or Docker. It has 5 stars and was created 2026-09-24.",
  "chunks": [
    {
      "index": 0,
      "heading_path": [],
      "est_tokens": 83,
      "text": "proxy-scraper is a self-hostable proxy scraper and verifier that pulls candidate proxies from 700+ sources, then runs honeypot, injected-script and TLS checks, and can serve a rotating proxy plus an MCP server for AI agents. Open source: yes (MIT); self-hostable via pipx, pip or Docker. It has 5 stars and was created 2026-09-24."
    },
    {
      "index": 1,
      "heading_path": [
        null,
        "What it does"
      ],
      "est_tokens": 253,
      "text": "self-hostable proxy scraper and verifier that pulls candidate proxies from 700+ sources, then runs honeypot, injected-script and TLS checks, and can serve a rotating proxy plus an MCP server for AI agents. Open source: yes (MIT); self-hostable via pipx, pip or Docker. It has 5 stars and was created 2026-09-24.\n\nproxy-scraper is two stages: collection and verification. It gathers HTTP, SOCKS4 and SOCKS5 candidates from 700+ public sources, then subjects each to a verification battery — a honeypot confirmation fetch, content-tampering/injected-script detection, TLS verification for HTTPS, real-IP validation, and two independent page fetches before a proxy is trusted. Survivors can be served two ways: `--serve` stands up a local rotating proxy on port 8899 (weighted, round-robin or fastest-only), and `proxy-scraper-mcp` exposes tools (`get_proxies`, `check_proxies`, `fetch_url`) so an AI agent can route requests through vetted proxies. Open source: yes (MIT); self-hostable via pipx, pip or Docker."
    },
    {
      "index": 2,
      "heading_path": [
        null,
        "Provenance"
      ],
      "est_tokens": 256,
      "text": "proxy is trusted. Survivors can be served two ways: `--serve` stands up a local rotating proxy on port 8899 (weighted, round-robin or fastest-only), and `proxy-scraper-mcp` exposes tools (`get_proxies`, `check_proxies`, `fetch_url`) so an AI agent can route requests through vetted proxies. Open source: yes (MIT); self-hostable via pipx, pip or Docker.\n\n- MIT (OSI-open); 5 stars, created 2026-09-24, last push 2026-09-26; pulls proxies from 700+ sources and verifies via honeypot, injected-script and TLS checks; ships a rotating proxy server and an MCP server; installs via pipx/pip/Docker (WebFetch 2026-09-26).\n- Surfaced via studio discovery in the 2026-09-26 pass.\n- Anti-hype note — load-bearing: the README's 'one in five working proxies injected a script' figure is asserted with no backing dataset in the repo; it is the tool's own unverified claim and is not stated as fact anywhere in this entry.\n- Curated from the GTM Stacker signal registry (2026-09-26 pass); license/facts independently verified 2026-09-26."
    },
    {
      "index": 3,
      "heading_path": [
        null,
        "Why it matters for a GTM stack"
      ],
      "est_tokens": 293,
      "text": "README's 'one in five working proxies injected a script' figure is asserted with no backing dataset in the repo; it is the tool's own unverified claim and is not stated as fact anywhere in this entry. - Curated from the GTM Stacker signal registry (2026-09-26 pass); license/facts independently verified 2026-09-26.\n\nScraping and enrichment pipelines usually rent proxies from a managed vendor, which means recurring cost and a dependency you do not control. proxy-scraper's angle is owned proxy infrastructure: collect candidates yourself, verify them yourself, and rotate them through a local server you run — with an MCP endpoint so an agent-driven enrichment step can fetch through vetted proxies directly. For a team building agent-driven scraping or prospecting enrichment, that is the difference between a proxy line-item and infrastructure you own. Open source: yes (MIT); pricing: free; self-hostable. The honest read: at 5 stars it is early, the injected-script ratio it cites is unverified vendor framing, and public proxies carry trust and legal risk no verifier fully removes — evaluate the verification mechanics on your own traffic before relying on them."
    }
  ],
  "alternates": {
    "markdown": "https://gtmstacker.com/registry/tool/proxy-scraper/index.md",
    "html": "https://gtmstacker.com/registry/tool/proxy-scraper/",
    "json": "https://gtmstacker.com/registry/tool/proxy-scraper/index.json",
    "server_json": "https://gtmstacker.com/registry/tool/proxy-scraper/server.json"
  },
  "jsonld": {
    "@context": "https://schema.org",
    "@graph": [
      {
        "@type": "WebSite",
        "@id": "https://gtmstacker.com/#website",
        "url": "https://gtmstacker.com/",
        "name": "GTM Stacker Agent Registry",
        "description": "A daily-updated, agent-native registry of open-source tool discoveries, tool updates, and curated news for the go-to-market / RevOps engineering niche. Machine-readable first: agents can discover, parse, page, and delta-sync it without scraping HTML.",
        "inLanguage": "en",
        "publisher": {
          "@id": "https://gtmstacker.com/#organization"
        }
      },
      {
        "@type": "Organization",
        "@id": "https://gtmstacker.com/#organization",
        "name": "GTM Stacker",
        "url": "https://gtmstacker.com",
        "description": "The growth-systems practice of Theo Popov: AI-native enrichment, outbound, content engines and internal tooling for startups and venture programs. Its agent-native media property, the GTM Stacker Agent Registry, maintains a daily-updated catalog of open-source go-to-market and RevOps tools that both people and AI engines can discover, compare, and cite.",
        "foundingDate": "2024-08",
        "knowsAbout": [
          "go-to-market engineering",
          "RevOps",
          "sales automation",
          "marketing operations",
          "open-source software",
          "AI agents"
        ],
        "founder": {
          "@type": "Person",
          "@id": "https://gtmstacker.com/#founder",
          "name": "Theo Popov",
          "jobTitle": "Growth Operations & GTM Systems",
          "url": "https://gtmstacker.com/about/",
          "sameAs": [
            "https://www.linkedin.com/in/theo-popov",
            "https://x.com/Theo_Popov",
            "https://github.com/theopopov"
          ],
          "worksFor": {
            "@id": "https://gtmstacker.com/#organization"
          }
        },
        "sameAs": [
          "https://www.linkedin.com/company/gtmstacker",
          "https://www.youtube.com/@gtmstacker",
          "https://www.instagram.com/gtmstacker/",
          "https://www.tiktok.com/@gtmstacker"
        ],
        "mainEntityOfPage": "https://gtmstacker.com/registry/about/"
      },
      {
        "@type": "SoftwareApplication",
        "@id": "https://gtmstacker.com/registry/tool/proxy-scraper/#software",
        "name": "proxy-scraper",
        "identifier": "maximilianfeix/proxy-scraper",
        "description": "proxy-scraper is a self-hostable proxy scraper and verifier that pulls candidate proxies from 700+ sources, then runs honeypot, injected-script and TLS checks, and can serve a rotating proxy plus an MCP server. Open source: yes (MIT); self-hostable via pipx, pip or Docker. 5 stars, created 2026-09-24.",
        "applicationCategory": "DeveloperApplication",
        "url": "https://gtmstacker.com/registry/tool/proxy-scraper/",
        "datePublished": "2026-09-26T00:00:00Z",
        "dateModified": "2026-09-26T00:00:00Z",
        "isPartOf": {
          "@id": "https://gtmstacker.com/#website"
        },
        "license": "https://spdx.org/licenses/MIT.html",
        "codeRepository": "https://github.com/maximilianfeix/proxy-scraper",
        "keywords": "data-scraping, mcp-agents, self-hostable, proxies, verification, rotating-proxy",
        "author": {
          "@type": "Organization",
          "name": "maximilianfeix",
          "url": "https://github.com/maximilianfeix",
          "sameAs": [
            "https://github.com/maximilianfeix/proxy-scraper"
          ]
        },
        "offers": {
          "@type": "Offer",
          "price": 0,
          "priceCurrency": "USD"
        },
        "isSimilarTo": [
          {
            "@type": "SoftwareApplication",
            "name": "DeepScrape",
            "url": "https://gtmstacker.com/registry/tool/deepscrape/",
            "applicationCategory": "DeveloperApplication",
            "offers": {
              "@type": "Offer",
              "price": 0,
              "priceCurrency": "USD"
            }
          },
          {
            "@type": "SoftwareApplication",
            "name": "Scrapy",
            "url": "https://gtmstacker.com/registry/tool/scrapy/",
            "applicationCategory": "DeveloperApplication",
            "offers": {
              "@type": "Offer",
              "price": 0,
              "priceCurrency": "USD"
            }
          }
        ]
      },
      {
        "@type": "BreadcrumbList",
        "@id": "https://gtmstacker.com/registry/tool/proxy-scraper/#breadcrumb",
        "itemListElement": [
          {
            "@type": "ListItem",
            "position": 1,
            "name": "GTM Stacker Registry",
            "item": "https://gtmstacker.com/registry/"
          },
          {
            "@type": "ListItem",
            "position": 2,
            "name": "Data Scraping",
            "item": "https://gtmstacker.com/registry/category/data-scraping/"
          },
          {
            "@type": "ListItem",
            "position": 3,
            "name": "proxy-scraper",
            "item": "https://gtmstacker.com/registry/tool/proxy-scraper/"
          }
        ]
      }
    ]
  },
  "tool": {
    "name": "maximilianfeix/proxy-scraper",
    "repository": {
      "url": "https://github.com/maximilianfeix/proxy-scraper",
      "source": "github"
    }
  }
}
