textsnap turns any image, screenshot or webpage into plaintext on-device via ONNX OCR, running fully offline after one model download — no GPU or API keys.
Yes, textsnap is open source under the MIT license.
textsnap is free to use.
Yes, textsnap can be self-hosted (the source is available under the MIT license).
textsnap converts any image, screenshot, or webpage into plaintext locally via OCR, and after the first model download it runs fully offline. Open source: yes (MIT for the project; the models are Apache-2.0); self-hostable with a pip install on the ONNX runtime — no GPU, no cloud, and no API keys. It is free, has ~186 stars, and is maintained by kouhxp.
textsnap is an on-device OCR primitive aimed at builders who need to turn visual content into text without calling a cloud OCR service. You point it at an image, a screenshot, or a webpage and it returns plaintext, running the model through the ONNX runtime on the local CPU — no GPU required and no API keys to manage. After the initial model download it works entirely offline, so nothing you OCR ever leaves the machine. It installs with pip, which makes it easy to drop into an agent or document-ingestion pipeline as the step that feeds text downstream. Open source: yes (MIT project; Apache-2.0 models); self-hostable; free.
GTM pipelines constantly hit text that is trapped in images — a screenshot of a pricing page, a scanned contract, a chart in a deck, a competitor's site rendered as an image. textsnap turns that into plaintext an agent can read, and it does it locally, which matters when the source contains confidential prospect or deal information you would rather not send to a third-party OCR API. The honest read: OCR accuracy on messy real-world inputs — skew, low resolution, dense tables, handwriting — is not benchmarked here, so test it on your actual documents; and while it runs offline afterward, the first run needs network access to fetch the model.