# Strata

> Updated 2026-10-03 · type: tool · category: ai-infrastructure · status: active · rev 1

Strata is an open-source inference engine that runs large open models on one consumer GPU, with an OpenAI and Anthropic-compatible API on localhost.

- Open source: yes (MIT)
- Self-hostable: yes
- Pricing model: free
- Best for: A team that wants to run large open models locally on a single consumer GPU and call them through an OpenAI or Anthropic-compatible API on localhost, so existing agents and tooling point at their own box instead of a hosted provider.
- Not for: A team with no suitable GPU (it needs an NVIDIA or AMD card with 12 GB or more), or one that wants a mature, battle-tested runtime rather than an early v0.1.x project, or that would rather call a managed API than operate inference itself.
- Last verified: 2026-10-03

- **Canonical:** https://gtmstacker.com/registry/tool/strata/
- **Source:** [Niko1221 · GitHub](https://github.com/Niko1221/Strata)
- **Tags:** ai-infrastructure, mcp-agents, inference, local-models, self-hostable, openai-compatible
- **Repository:** https://github.com/Niko1221/Strata

## Is Strata open source?

Yes, Strata is open source under the MIT license.

## How much does Strata cost?

Strata is free to use.

## Can I self-host Strata?

Yes, Strata can be self-hosted (the source is available under the MIT license).

## Alternatives & related

- [Magnitude](https://gtmstacker.com/registry/tool/magnitude/)


---

Strata is an open-source local inference engine: it runs large open models on a single consumer GPU and exposes an OpenAI and Anthropic-compatible API on localhost, so agents and tools you already have can point at your own machine instead of a hosted provider. Open source: yes (MIT); self-hostable; free OSS. It has ~8.2k stars, ~628 commits, and the maker demonstrates a 125-billion-parameter model running on a gaming PC with an NVIDIA or AMD card of 12 GB or more.

## What it does

Strata is about running big models on hardware you own. The pitch is a one-click install for Windows or Linux that stands up an inference engine on your GPU and serves it through an OpenAI/Anthropic-compatible endpoint on localhost, with optional image input. Because the API mirrors the hosted providers, you can repoint an existing agent at it without rewriting the integration. The demonstration that draws attention is a 125-billion-parameter model on a single gaming PC, achieved with heavy quantisation (the README shows IQ3_S at 128K context). Open source: yes (MIT); self-hostable; free OSS. It needs an NVIDIA or AMD GPU with 12 GB or more.

## Provenance

- MIT per repo (license read from GitHub metadata, not guessed); ~8,158 stars; ~628 commits; active (pushed 2026-10-03) (github.com/Niko1221/Strata, verified 2026-10-03).
- README: "Run a 125-billion-parameter AI model on your own gaming PC"; NVIDIA or AMD card with 12 GB or more; Windows or Linux; one-click install; OpenAI/Anthropic-compatible API on localhost; optional image input.
- Surfaced in the GTM Stacker X/Twitter signal report (2026-10-03 pass); license/facts independently verified 2026-10-03.
- The 125B-on-a-consumer-GPU demo uses aggressive quantisation (IQ3_S) and is the maker's own (vendor-claim); not independently benchmarked here.

## Why it matters for a GTM stack

Running your own models is where agent economics stop being a hosted-API line item, and the barrier has always been hardware. Strata's bet is that heavy quantisation plus a one-click install lets a single consumer GPU serve a model large enough to be useful, behind the same OpenAI/Anthropic API your agents already speak. For a GTM team that wants prospecting, enrichment, or support agents running on owned infrastructure with no per-token bill, that is worth a trial on a spare box. The honest read: the headline number leans on IQ3_S quantisation, which costs output quality, and it is an early v0.1.x project, so benchmark the quality on your own tasks and treat the "125B on a gaming PC" line as a demo, not a guarantee. You also need a 12 GB-plus NVIDIA or AMD card to start.
