{
  "$schema_doc": "https://gtmstacker.com/registry/schema/entry.schema.json",
  "stability": "emerging",
  "generator": "agentic-media-registry",
  "generated_at": "2026-09-23T00:00:00Z",
  "id": "com.gtmstacker.registry/tool/tenderness",
  "type": "tool",
  "slug": "tenderness",
  "canonical_url": "https://gtmstacker.com/registry/tool/tenderness/",
  "title": "Tenderness",
  "description": "Apache-2.0 Python library that renders synthetic, deterministic documents from text and images via Cairo/Pango to produce labeled datasets for training vision-language and OCR models. It is a rendering library, not a hosted generator app — you call it in code to emit page images with known ground-truth, so the data-prep layer for a document-extraction model is reproducible instead of hand-collected. Early and small (7 stars), for teams needing controlled training data, not a turnkey product.",
  "category": "ai-infrastructure",
  "tags": [
    "ai-infrastructure",
    "data-scraping",
    "self-hostable",
    "ocr",
    "synthetic-data",
    "python-library"
  ],
  "status": "active",
  "revision": 1,
  "content_hash": "4f3f4b4de02044bd0cba8ccd9f37481a3dd74d7aa04c3110a4e8209efe5f5fd4",
  "date_published": "2026-09-23T00:00:00Z",
  "date_modified": "2026-09-23T00:00:00Z",
  "source": {
    "name": "github · paperchase-labs/tenderness",
    "url": "https://github.com/paperchase-labs/tenderness"
  },
  "license": "Apache-2.0",
  "one_liner": "Tenderness renders synthetic labeled documents via Cairo/Pango to build training datasets for VLM and OCR models.",
  "open_source": "yes",
  "self_hostable": "yes",
  "pricing_model": "free",
  "who_its_for": "An ML or data engineer building a document-extraction, OCR or VLM pipeline who needs deterministic, labeled page images as training data and wants to generate them in code rather than scrape and annotate real documents.",
  "aliases": [
    "tenderness",
    "paperchase-labs/tenderness"
  ],
  "alternatives": [],
  "secondary_categories": [
    "data-scraping",
    "prospecting-enrichment"
  ],
  "last_verified": "2026-09-23",
  "evidence": {
    "claim_type": "mixed",
    "source_id": "https://github.com/paperchase-labs/tenderness",
    "note": "Apache-2.0 confirmed in repo metadata; 7 stars, last push 2026-08-04; installs via `pip install tenderness`; renders text/images to documents through Cairo/Pango for VLM/OCR training data (WebFetch 2026-09-23, github.com/paperchase-labs/tenderness). Framing correction: it is a rendering library you import, not a general document-generator application."
  },
  "caveats": "Small and early — 7 stars, last pushed 2026-08-04 — so treat it as a library to evaluate, not a maintained platform. It is deliberately narrow: it renders documents deterministically for dataset creation; it does not train models, run OCR, or extract fields for you.",
  "lead": "Apache-2.0 Python library that renders synthetic, deterministic documents from text and images via Cairo/Pango to produce labeled datasets for training vision-language and OCR models. It is a rendering library, not a hosted generator app — you call it in code to emit page images with known ground-truth, so the data-prep…",
  "chunks": [
    {
      "index": 0,
      "heading_path": [],
      "est_tokens": 128,
      "text": "Apache-2.0 Python library that renders synthetic, deterministic documents from text and images via Cairo/Pango to produce labeled datasets for training vision-language and OCR models. It is a rendering library, not a hosted generator app — you call it in code to emit page images with known ground-truth, so the data-prep layer for a document-extraction model is reproducible instead of hand-collected. Early and small (7 stars), it targets teams that need controlled training data rather than a turnkey product."
    },
    {
      "index": 1,
      "heading_path": [
        null,
        "Provenance"
      ],
      "est_tokens": 218,
      "text": "library, not a hosted generator app — you call it in code to emit page images with known ground-truth, so the data-prep layer for a document-extraction model is reproducible instead of hand-collected. Early and small (7 stars), it targets teams that need controlled training data rather than a turnkey product.\n\n- Apache-2.0 (OSI-open, confirmed in repo metadata); 7 stars, last push 2026-08-04; installs via `pip install tenderness`; renders text and images to documents through Cairo/Pango to generate labeled datasets for VLM/OCR training (WebFetch 2026-09-23, github.com/paperchase-labs/tenderness).\n- Framing correction applied: this is a rendering library, not a general-purpose \"generator\" application — it is imported and driven from your own code.\n- Curated from the GTM Stacker signal registry (2026-09-23 pass); license/facts independently verified 2026-09-23."
    },
    {
      "index": 2,
      "heading_path": [
        null,
        "Why it matters for a GTM stack"
      ],
      "est_tokens": 290,
      "text": "through Cairo/Pango to generate labeled datasets for VLM/OCR training (WebFetch 2026-09-23, github.com/paperchase-labs/tenderness). - Framing correction applied: this is a rendering library, not a general-purpose \"generator\" application — it is imported and driven from your own code. - Curated from the GTM Stacker signal registry (2026-09-23 pass); license/facts independently verified 2026-09-23.\n\nDocument extraction is a common enrichment primitive — parse invoices, contracts, forms, scanned lead lists — and the bottleneck is almost never the model, it is labeled data that matches your document shapes. Tenderness sits underneath that: it renders synthetic pages with known ground-truth deterministically, so a team can generate the exact training set a custom VLM/OCR extractor needs instead of scraping and hand-annotating real documents. For a RevOps or data team building its own document-enrichment pipeline, that is the reproducible data-prep layer. The honest read: it is small (7 stars, last pushed early August 2026) and deliberately narrow — a library for making training data, not a system that trains, reads or extracts anything on its own."
    }
  ],
  "alternates": {
    "markdown": "https://gtmstacker.com/registry/tool/tenderness/index.md",
    "html": "https://gtmstacker.com/registry/tool/tenderness/",
    "json": "https://gtmstacker.com/registry/tool/tenderness/index.json",
    "server_json": "https://gtmstacker.com/registry/tool/tenderness/server.json"
  },
  "jsonld": {
    "@context": "https://schema.org",
    "@graph": [
      {
        "@type": "WebSite",
        "@id": "https://gtmstacker.com/#website",
        "url": "https://gtmstacker.com/",
        "name": "GTM Stacker Agent Registry",
        "description": "A daily-updated, agent-native registry of open-source tool discoveries, tool updates, and curated news for the go-to-market / RevOps engineering niche. Machine-readable first: agents can discover, parse, page, and delta-sync it without scraping HTML.",
        "inLanguage": "en",
        "publisher": {
          "@id": "https://gtmstacker.com/#organization"
        }
      },
      {
        "@type": "Organization",
        "@id": "https://gtmstacker.com/#organization",
        "name": "GTM Stacker",
        "url": "https://gtmstacker.com",
        "description": "The growth-systems practice of Theo Popov: AI-native enrichment, outbound, content engines and internal tooling for startups and venture programs. Its agent-native media property, the GTM Stacker Agent Registry, maintains a daily-updated catalog of open-source go-to-market and RevOps tools that both people and AI engines can discover, compare, and cite.",
        "foundingDate": "2024-08",
        "knowsAbout": [
          "go-to-market engineering",
          "RevOps",
          "sales automation",
          "marketing operations",
          "open-source software",
          "AI agents"
        ],
        "founder": {
          "@type": "Person",
          "@id": "https://gtmstacker.com/#founder",
          "name": "Theo Popov",
          "jobTitle": "Growth Operations & GTM Systems",
          "url": "https://gtmstacker.com/about/",
          "sameAs": [
            "https://www.linkedin.com/in/theo-popov",
            "https://x.com/Theo_Popov",
            "https://github.com/theopopov"
          ],
          "worksFor": {
            "@id": "https://gtmstacker.com/#organization"
          }
        },
        "sameAs": [
          "https://www.linkedin.com/company/gtmstacker",
          "https://www.youtube.com/@gtmstacker",
          "https://www.instagram.com/gtmstacker/",
          "https://www.tiktok.com/@gtmstacker"
        ],
        "mainEntityOfPage": "https://gtmstacker.com/registry/about/"
      },
      {
        "@type": "SoftwareApplication",
        "@id": "https://gtmstacker.com/registry/tool/tenderness/#software",
        "name": "Tenderness",
        "identifier": "io.github.paperchase-labs/tenderness",
        "description": "Apache-2.0 Python library that renders synthetic, deterministic documents from text and images via Cairo/Pango to produce labeled datasets for training vision-language and OCR models. It is a rendering library, not a hosted generator app — you call it in code to emit page images with known ground-truth, so the data-prep layer for a document-extraction model is reproducible instead of hand-collected. Early and small (7 stars), for teams needing controlled training data, not a turnkey product.",
        "applicationCategory": "DeveloperApplication",
        "url": "https://gtmstacker.com/registry/tool/tenderness/",
        "datePublished": "2026-09-23T00:00:00Z",
        "dateModified": "2026-09-23T00:00:00Z",
        "isPartOf": {
          "@id": "https://gtmstacker.com/#website"
        },
        "license": "https://spdx.org/licenses/Apache-2.0.html",
        "codeRepository": "https://github.com/paperchase-labs/tenderness",
        "keywords": "ai-infrastructure, data-scraping, prospecting-enrichment, self-hostable, ocr, synthetic-data, python-library",
        "author": {
          "@type": "Organization",
          "name": "paperchase-labs",
          "url": "https://github.com/paperchase-labs",
          "sameAs": [
            "https://github.com/paperchase-labs/tenderness"
          ]
        },
        "offers": {
          "@type": "Offer",
          "price": 0,
          "priceCurrency": "USD"
        }
      },
      {
        "@type": "BreadcrumbList",
        "@id": "https://gtmstacker.com/registry/tool/tenderness/#breadcrumb",
        "itemListElement": [
          {
            "@type": "ListItem",
            "position": 1,
            "name": "GTM Stacker Registry",
            "item": "https://gtmstacker.com/registry/"
          },
          {
            "@type": "ListItem",
            "position": 2,
            "name": "AI Infrastructure",
            "item": "https://gtmstacker.com/registry/category/ai-infrastructure/"
          },
          {
            "@type": "ListItem",
            "position": 3,
            "name": "Tenderness",
            "item": "https://gtmstacker.com/registry/tool/tenderness/"
          }
        ]
      }
    ]
  },
  "tool": {
    "name": "io.github.paperchase-labs/tenderness",
    "repository": {
      "url": "https://github.com/paperchase-labs/tenderness",
      "source": "github"
    }
  }
}
