# Jeeves

> Updated 2026-09-29 · type: tool · category: ai-infrastructure · status: active · rev 1

Jeeves is a Qwen3.5-9B decision classifier that returns calibrated yes/no/rating calls together with the reasoning trace behind each verdict.

- Open source: yes (MIT)
- Self-hostable: yes
- Pricing model: free
- Best for: A team that wants a self-hosted decision classifier whose calibrated yes/no or rating output comes with a visible reasoning trace — so a human can audit why a call was made — running locally via CUDA or the Python SDK.
- Not for: A team on lightweight hardware that cannot host a 9B model, or one that needs sub-30ms inline decisions rather than reasoned, auditable judgments.
- Last verified: 2026-09-29

- **Canonical:** https://gtmstacker.com/registry/tool/jeeves/
- **Source:** [PostHog · GitHub](https://github.com/PostHog/jeeves)
- **Tags:** ai-infrastructure, decision-model, classification, reasoning, self-hostable, local-inference
- **Repository:** https://github.com/PostHog/jeeves

## Is Jeeves open source?

Yes, Jeeves is open source under the MIT license.

## How much does Jeeves cost?

Jeeves is free to use.

## Can I self-host Jeeves?

Yes, Jeeves can be self-hosted (the source is available under the MIT license).

## Alternatives & related

- [Laya](https://gtmstacker.com/registry/tool/laya/)
- [Jevpipe](https://gtmstacker.com/registry/tool/jevpipe/)
- [Jeff](https://gtmstacker.com/registry/tool/jeff/)


---

Jeeves is a reasoning-augmented decision classifier built on Qwen3.5-9B that produces calibrated yes/no and rating calls while exposing the reasoning behind each one. Open source: yes (MIT); self-hostable by downloading the weights and running local CUDA inference or using the Python SDK. It is free, has ~211 stars, and is maintained by PostHog — a well-known analytics company, which is a useful credibility signal for the project's provenance.

## What it does

Jeeves takes a decision prompt and returns a structured verdict — yes/no or a rating — that the project describes as calibrated, and crucially it surfaces the reasoning trace that led to the verdict rather than emitting a bare label. That transparency is the differentiator: where a small classifier gives you an answer, Jeeves gives you an answer you can audit. It is built on Qwen3.5-9B, so it is a larger model than the sub-2B single-purpose decision models; you self-host it by downloading the weights and running CUDA inference or calling it through the Python SDK. Open source: yes (MIT); self-hostable; free.

## Provenance

- MIT per repo; ~211 stars.
- Reasoning-augmented Qwen3.5-9B decision classifier; calibrated yes/no/rating output with visible reasoning.
- Self-hostable: download weights, run local CUDA inference or use the Python SDK (PostHog/jeeves, verified 2026-09-29).
- Maker is PostHog, an established analytics company — a provenance/credibility signal.
- Surfaced via the GTM Stacker studio daily pull (2026-09-29 pass); license/facts verified against the primary repo 2026-09-29.

## Why it matters for a GTM stack

Many GTM decisions are judgment calls that later need defending — is this lead a fit, does this reply signal intent, should this account be escalated. A bare classifier label is hard to trust or debug; Jeeves pairs the calibrated verdict with the reasoning behind it, so an operator can see why the model decided as it did and catch systematic errors before they compound across a pipeline. The honest read: "calibrated" is the vendor's own claim and should be checked against your labeled outcomes, and a 9B model is materially heavier to self-host than the tiny decision models — so choose Jeeves when the auditability of each call is worth the extra GPU footprint, and a smaller model when you only need a fast inline label.
