# ML Drift

> Updated 2026-10-10 · type: tool · category: ai-infrastructure · status: active · rev 1

ML Drift is Google's open-source GPU engine that runs on-device AI inference on mobile, web, and desktop without a server round-trip.

- Open source: yes (Apache-2.0)
- Self-hostable: yes
- Pricing model: free
- Best for: A team shipping AI features into a mobile app, browser extension, or desktop client who wants inference to run on the user's own device GPU instead of a metered API call, for latency, cost, or privacy reasons.
- Not for: Teams running large models server-side where device-GPU constraints don't apply — ML Drift targets on-device inference, not datacenter-scale serving.
- Last verified: 2026-10-10

- **Canonical:** https://gtmstacker.com/registry/tool/ml-drift/
- **Source:** [github · google-ai-edge/ml-drift](https://github.com/google-ai-edge/ml-drift)
- **Tags:** ai-infrastructure, self-hostable, on-device, inference, open-source
- **Repository:** https://github.com/google-ai-edge/ml-drift

## Is ML Drift open source?

Yes, ML Drift is open source under the Apache-2.0 license.

## How much does ML Drift cost?

ML Drift is free to use.

## Can I self-host ML Drift?

Yes, ML Drift can be self-hosted (the source is available under the Apache-2.0 license).


---

Google's open-source, cross-platform GPU-accelerated inference engine for on-device machine learning — the same GPU backend that has powered TensorFlow Lite/LiteRT inside Google products, now published as a standalone Apache-2.0 repo. Runs AI/ML workloads, including large generative models, on mobile, web, and desktop GPUs via OpenCL, Metal, WebGPU, and OpenGL ES 3.1+.

## Provenance

- Apache-2.0, 117 stars, 12 forks — independently WebFetch-verified 2026-10-10 (github.com/google-ai-edge/ml-drift).
- Surfaced via the 2026-10-10 OSS X pull (Google AI Edge team announcement thread).
- Curated from the GTM Stacker signal registry (2026-10-10 pass); license independently verified 2026-10-10.

## Why it matters for a GTM stack

Shipping an AI feature into a consumer-facing app usually means a metered API call per inference, which adds latency, cost, and a data round-trip a privacy-sensitive user or team may not want. ML Drift's bet is that the device GPU already sitting in the user's phone or browser can do real inference work for free, if the engine abstracts away the OpenCL/Metal/WebGPU/OpenGL differences. The honest read: this is Google's internal GPU backend newly published standalone, so the headline performance numbers (up to 40% latency cut, up to 2 seconds faster) are Google's own production results on Google's own products, not independently re-verified on a general third-party app — worth prototyping on your actual model and device mix before assuming the same gains.
