{
  "$schema_doc": "https://gtmstacker.com/registry/schema/entry.schema.json",
  "stability": "emerging",
  "generator": "agentic-media-registry",
  "generated_at": "2026-09-29T00:00:00Z",
  "id": "com.gtmstacker.registry/tool/jeff",
  "type": "tool",
  "slug": "jeff",
  "canonical_url": "https://gtmstacker.com/registry/tool/jeff/",
  "title": "Jeff",
  "description": "Jeff is a family of small (0.8B–2B) zero-shot decision and classification models you embed directly in code for routing, moderation, and intent calls at ~22–28ms latency. Open source: yes (MIT code + Apache-2.0 model weights); self-hostable (local jeff-serve on CUDA or Apple MLX, weights pulled from HuggingFace, no external API). Free; ~952 stars; maker firelex.",
  "category": "ai-infrastructure",
  "tags": [
    "ai-infrastructure",
    "decision-model",
    "classification",
    "zero-shot",
    "local-inference",
    "self-hostable"
  ],
  "status": "active",
  "revision": 1,
  "content_hash": "93e8b1464746d2f09664f35a5fc46ac89edd582fbacca169ffa2db01d95cb4d4",
  "date_published": "2026-09-29T00:00:00Z",
  "date_modified": "2026-09-29T00:00:00Z",
  "source": {
    "name": "firelex · GitHub",
    "url": "https://github.com/firelex/jeff"
  },
  "license": "MIT",
  "one_liner": "Jeff embeds 0.8B–2B zero-shot decision models in code, returning routing, moderation and intent calls in ~22–28ms with no external API.",
  "open_source": "yes",
  "self_hostable": "yes",
  "pricing_model": "free",
  "who_its_for": "A builder who needs fast, cheap decision or classification calls — route, moderate, classify intent — inline in application code, running the model locally on CUDA or Apple MLX with no per-call API cost or network round-trip.",
  "who_its_not_for": "A team that needs a general-purpose reasoning LLM, long-form generation, or a large-context model — these are small single-purpose decision models, not chat models.",
  "aliases": [
    "jeff",
    "firelex/jeff"
  ],
  "alternatives": [
    "laya",
    "jevpipe"
  ],
  "secondary_categories": [],
  "last_verified": "2026-09-29",
  "evidence": {
    "claim_type": "vendor-claim",
    "source_id": "https://github.com/firelex/jeff",
    "note": "MIT for code, Apache-2.0 for model weights, per repo; ~952 stars. 0.8B–2B zero-shot decision/classification models; vendor reports ~22–28ms latency. Self-hostable via local `jeff-serve` on CUDA or Apple MLX; weights pulled from HuggingFace; no external API (firelex/jeff, verified 2026-09-29)."
  },
  "caveats": "Verified from the primary repo (2026-09-29); no independent benchmarking here. The ~22–28ms latency is the vendor's own figure (vendor-claim) — it will depend on your hardware, model size (0.8B vs 2B), and batch settings; measure on your target CPU/GPU. Dual license: MIT for the code, Apache-2.0 for the weights.",
  "lead": "Jeff is a family of small (0.8B–2B) zero-shot decision and classification models designed to be embedded directly in application code — routing, moderation, and intent calls returned in roughly 22–28ms. Open source: yes (MIT for the code, Apache-2.0 for the model weights); self-hostable via a local `jeff-serve` on CUDA or…",
  "chunks": [
    {
      "index": 0,
      "heading_path": [],
      "est_tokens": 113,
      "text": "Jeff is a family of small (0.8B–2B) zero-shot decision and classification models designed to be embedded directly in application code — routing, moderation, and intent calls returned in roughly 22–28ms. Open source: yes (MIT for the code, Apache-2.0 for the model weights); self-hostable via a local `jeff-serve` on CUDA or Apple MLX, with weights pulled from HuggingFace and no external API. It is free, has ~952 stars, and is maintained by firelex."
    },
    {
      "index": 1,
      "heading_path": [
        null,
        "What it does"
      ],
      "est_tokens": 240,
      "text": "moderation, and intent calls returned in roughly 22–28ms. Open source: yes (MIT for the code, Apache-2.0 for the model weights); self-hostable via a local `jeff-serve` on CUDA or Apple MLX, with weights pulled from HuggingFace and no external API. It is free, has ~952 stars, and is maintained by firelex.\n\nJeff replaces a hosted-LLM API call for the narrow class of decisions that dominate agent and application plumbing — \"which route does this go to,\" \"is this content allowed,\" \"what is the user's intent.\" Instead of paying per-token latency and cost to a remote model, you run a 0.8B–2B model locally and get a typed decision back in the tens of milliseconds. The models are zero-shot, so you describe the labels or the decision rather than fine-tuning. You serve them yourself with `jeff-serve` on CUDA or Apple MLX; the weights come from HuggingFace and nothing leaves the machine. Open source: yes (MIT code + Apache-2.0 weights); self-hostable; free."
    },
    {
      "index": 2,
      "heading_path": [
        null,
        "Provenance"
      ],
      "est_tokens": 192,
      "text": "the tens of milliseconds. The models are zero-shot, so you describe the labels or the decision rather than fine-tuning. You serve them yourself with `jeff-serve` on CUDA or Apple MLX; the weights come from HuggingFace and nothing leaves the machine. Open source: yes (MIT code + Apache-2.0 weights); self-hostable; free.\n\n- MIT license for the code, Apache-2.0 for the model weights, per repo; ~952 stars.\n- 0.8B–2B zero-shot decision/classification models; vendor-reported latency ~22–28ms.\n- Self-hostable via local `jeff-serve` on CUDA or Apple MLX; weights pulled from HuggingFace; no external API (firelex/jeff, verified 2026-09-29).\n- Surfaced via the GTM Stacker studio daily pull (2026-09-29 pass); license/facts verified against the primary repo 2026-09-29."
    },
    {
      "index": 3,
      "heading_path": [
        null,
        "Why it matters for a GTM stack"
      ],
      "est_tokens": 286,
      "text": "per repo; ~952 stars. - 0.8B–2B zero-shot decision/classification models; vendor-reported latency ~22–28ms. - Self-hostable via local `jeff-serve` on CUDA or Apple MLX; weights pulled from HuggingFace; no external API (firelex/jeff, verified 2026-09-29). - Surfaced via the GTM Stacker studio daily pull (2026-09-29 pass); license/facts verified against the primary repo 2026-09-29.\n\nA GTM agent makes hundreds of small decisions per run — route this lead to the right sequence, flag this reply as out-of-office, classify this inbound intent — and paying a hosted LLM for each one is slow and expensive at volume. Jeff lets you push those decisions into a local model that answers in tens of milliseconds with no per-call cost and no data leaving your infrastructure, which matters when the input is prospect or customer text. The honest read: the ~22–28ms figure is the vendor's and depends heavily on your hardware and model size, and these are single-purpose decision models — not a substitute for a reasoning LLM — so scope them to the routing/moderation/intent tier and benchmark on your own machine before wiring them into the hot path."
    }
  ],
  "alternates": {
    "markdown": "https://gtmstacker.com/registry/tool/jeff/index.md",
    "html": "https://gtmstacker.com/registry/tool/jeff/",
    "json": "https://gtmstacker.com/registry/tool/jeff/index.json",
    "server_json": "https://gtmstacker.com/registry/tool/jeff/server.json"
  },
  "jsonld": {
    "@context": "https://schema.org",
    "@graph": [
      {
        "@type": "WebSite",
        "@id": "https://gtmstacker.com/#website",
        "url": "https://gtmstacker.com/",
        "name": "GTM Stacker Agent Registry",
        "description": "A daily-updated, agent-native registry of open-source tool discoveries, tool updates, and curated news for the go-to-market / RevOps engineering niche. Machine-readable first: agents can discover, parse, page, and delta-sync it without scraping HTML.",
        "inLanguage": "en",
        "publisher": {
          "@id": "https://gtmstacker.com/#organization"
        }
      },
      {
        "@type": "Organization",
        "@id": "https://gtmstacker.com/#organization",
        "name": "GTM Stacker",
        "url": "https://gtmstacker.com",
        "description": "The growth-systems practice of Theo Popov: AI-native enrichment, outbound, content engines and internal tooling for startups and venture programs. Its agent-native media property, the GTM Stacker Agent Registry, maintains a daily-updated catalog of open-source go-to-market and RevOps tools that both people and AI engines can discover, compare, and cite.",
        "foundingDate": "2024-08",
        "knowsAbout": [
          "go-to-market engineering",
          "RevOps",
          "sales automation",
          "marketing operations",
          "open-source software",
          "AI agents"
        ],
        "founder": {
          "@type": "Person",
          "@id": "https://gtmstacker.com/#founder",
          "name": "Theo Popov",
          "jobTitle": "Growth Operations & GTM Systems",
          "url": "https://gtmstacker.com/about/",
          "sameAs": [
            "https://www.linkedin.com/in/theo-popov",
            "https://x.com/Theo_Popov",
            "https://github.com/theopopov"
          ],
          "worksFor": {
            "@id": "https://gtmstacker.com/#organization"
          }
        },
        "sameAs": [
          "https://www.linkedin.com/company/gtmstacker",
          "https://www.youtube.com/@gtmstacker",
          "https://www.instagram.com/gtmstacker/",
          "https://www.tiktok.com/@gtmstacker"
        ],
        "mainEntityOfPage": "https://gtmstacker.com/registry/about/"
      },
      {
        "@type": "SoftwareApplication",
        "@id": "https://gtmstacker.com/registry/tool/jeff/#software",
        "name": "Jeff",
        "identifier": "firelex/jeff",
        "description": "Jeff is a family of small (0.8B–2B) zero-shot decision and classification models you embed directly in code for routing, moderation, and intent calls at ~22–28ms latency. Open source: yes (MIT code + Apache-2.0 model weights); self-hostable (local jeff-serve on CUDA or Apple MLX, weights pulled from HuggingFace, no external API). Free; ~952 stars; maker firelex.",
        "applicationCategory": "DeveloperApplication",
        "url": "https://gtmstacker.com/registry/tool/jeff/",
        "datePublished": "2026-09-29T00:00:00Z",
        "dateModified": "2026-09-29T00:00:00Z",
        "isPartOf": {
          "@id": "https://gtmstacker.com/#website"
        },
        "license": "https://spdx.org/licenses/MIT.html",
        "codeRepository": "https://github.com/firelex/jeff",
        "keywords": "ai-infrastructure, decision-model, classification, zero-shot, local-inference, self-hostable",
        "author": {
          "@type": "Organization",
          "name": "firelex",
          "url": "https://github.com/firelex",
          "sameAs": [
            "https://github.com/firelex/jeff"
          ]
        },
        "offers": {
          "@type": "Offer",
          "price": 0,
          "priceCurrency": "USD"
        },
        "isSimilarTo": [
          {
            "@type": "SoftwareApplication",
            "name": "Laya",
            "url": "https://gtmstacker.com/registry/tool/laya/",
            "applicationCategory": "DeveloperApplication",
            "offers": {
              "@type": "Offer",
              "price": 0,
              "priceCurrency": "USD"
            }
          },
          {
            "@type": "SoftwareApplication",
            "name": "Jevpipe",
            "url": "https://gtmstacker.com/registry/tool/jevpipe/",
            "applicationCategory": "DeveloperApplication",
            "offers": {
              "@type": "Offer",
              "price": 0,
              "priceCurrency": "USD"
            }
          }
        ]
      },
      {
        "@type": "BreadcrumbList",
        "@id": "https://gtmstacker.com/registry/tool/jeff/#breadcrumb",
        "itemListElement": [
          {
            "@type": "ListItem",
            "position": 1,
            "name": "GTM Stacker Registry",
            "item": "https://gtmstacker.com/registry/"
          },
          {
            "@type": "ListItem",
            "position": 2,
            "name": "AI Infrastructure",
            "item": "https://gtmstacker.com/registry/category/ai-infrastructure/"
          },
          {
            "@type": "ListItem",
            "position": 3,
            "name": "Jeff",
            "item": "https://gtmstacker.com/registry/tool/jeff/"
          }
        ]
      }
    ]
  },
  "tool": {
    "name": "firelex/jeff",
    "repository": {
      "url": "https://github.com/firelex/jeff",
      "source": "github"
    }
  }
}
