# Jeff

> Updated 2026-09-29 · type: tool · category: ai-infrastructure · status: active · rev 1

Jeff embeds 0.8B–2B zero-shot decision models in code, returning routing, moderation and intent calls in ~22–28ms with no external API.

- Open source: yes (MIT)
- Self-hostable: yes
- Pricing model: free
- Best for: A builder who needs fast, cheap decision or classification calls — route, moderate, classify intent — inline in application code, running the model locally on CUDA or Apple MLX with no per-call API cost or network round-trip.
- Not for: A team that needs a general-purpose reasoning LLM, long-form generation, or a large-context model — these are small single-purpose decision models, not chat models.
- Last verified: 2026-09-29

- **Canonical:** https://gtmstacker.com/registry/tool/jeff/
- **Source:** [firelex · GitHub](https://github.com/firelex/jeff)
- **Tags:** ai-infrastructure, decision-model, classification, zero-shot, local-inference, self-hostable
- **Repository:** https://github.com/firelex/jeff

## Is Jeff open source?

Yes, Jeff is open source under the MIT license.

## How much does Jeff cost?

Jeff is free to use.

## Can I self-host Jeff?

Yes, Jeff can be self-hosted (the source is available under the MIT license).

## Alternatives & related

- [Laya](https://gtmstacker.com/registry/tool/laya/)
- [Jevpipe](https://gtmstacker.com/registry/tool/jevpipe/)


---

Jeff is a family of small (0.8B–2B) zero-shot decision and classification models designed to be embedded directly in application code — routing, moderation, and intent calls returned in roughly 22–28ms. Open source: yes (MIT for the code, Apache-2.0 for the model weights); self-hostable via a local `jeff-serve` on CUDA or Apple MLX, with weights pulled from HuggingFace and no external API. It is free, has ~952 stars, and is maintained by firelex.

## What it does

Jeff replaces a hosted-LLM API call for the narrow class of decisions that dominate agent and application plumbing — "which route does this go to," "is this content allowed," "what is the user's intent." Instead of paying per-token latency and cost to a remote model, you run a 0.8B–2B model locally and get a typed decision back in the tens of milliseconds. The models are zero-shot, so you describe the labels or the decision rather than fine-tuning. You serve them yourself with `jeff-serve` on CUDA or Apple MLX; the weights come from HuggingFace and nothing leaves the machine. Open source: yes (MIT code + Apache-2.0 weights); self-hostable; free.

## Provenance

- MIT license for the code, Apache-2.0 for the model weights, per repo; ~952 stars.
- 0.8B–2B zero-shot decision/classification models; vendor-reported latency ~22–28ms.
- Self-hostable via local `jeff-serve` on CUDA or Apple MLX; weights pulled from HuggingFace; no external API (firelex/jeff, verified 2026-09-29).
- Surfaced via the GTM Stacker studio daily pull (2026-09-29 pass); license/facts verified against the primary repo 2026-09-29.

## Why it matters for a GTM stack

A GTM agent makes hundreds of small decisions per run — route this lead to the right sequence, flag this reply as out-of-office, classify this inbound intent — and paying a hosted LLM for each one is slow and expensive at volume. Jeff lets you push those decisions into a local model that answers in tens of milliseconds with no per-call cost and no data leaving your infrastructure, which matters when the input is prospect or customer text. The honest read: the ~22–28ms figure is the vendor's and depends heavily on your hardware and model size, and these are single-purpose decision models — not a substitute for a reasoning LLM — so scope them to the routing/moderation/intent tier and benchmark on your own machine before wiring them into the hot path.
