# FreeToken

> Updated 2026-09-01 · type: tool · category: ai-infrastructure · status: active · rev 1

Edge-native Mixture-of-Experts inference engine (UC Berkeley) that runs frontier open-weight models on consumer hardware — e.g. a 35B model on an 8GB GPU — by…

- Open source: yes (Apache-2.0)
- Self-hostable: yes
- Pricing model: free
- Best for: Individuals with consumer/gaming hardware who want to run frontier-scale open-weight MoE models locally.
- Last verified: 2026-09-04

- **Canonical:** https://gtmstacker.com/registry/tool/freetoken/
- **Source:** [github · FlashML-org/FreeToken](https://github.com/FlashML-org/FreeToken)
- **Tags:** ai-infrastructure, local-inference, moe, self-hosting
- **Repository:** https://github.com/FlashML-org/FreeToken

## Is FreeToken open source?

Yes — Apache-2.0-licensed.

## How much does FreeToken cost?

It is free to use.

## Can I self-host FreeToken?

Yes — it can be self-hosted.


---

Edge-native Mixture-of-Experts inference engine (UC Berkeley) that runs frontier open-weight models on consumer hardware — e.g. a 35B model on an 8GB GPU — by treating GPU/CPU/host-memory as one bandwidth-adaptive platform; 2–4× faster than Ollama on MoE. Desktop app + Python CLI.

## Provenance

- Apache-2.0 verified (11.4k★, active; arXiv 2608.16157).
- Surfaced via the GTM Stacker X/Twitter signal reports (the "35 marketing repos" list + named drops); license independently WebFetch-verified 2026-09-03.
