Repo of the Day
soycaporal/ternlight
Published: Aug 30, 2026
Open repository ↗Contribute to soycaporal/ternlight development by creating an account on GitHub.
Summary
ternlight packages a small sentence-embedding model, BERT-style tokenizer, and inference engine into a single 5–7 MB WebAssembly file that runs on CPU. The weights are ternary (-1, 0, +1) and the Rust engine is compiled to WASM SIMD, so embeddings reduce to adds and subtracts rather than floating-point multiplies. It exposes a tiny JavaScript API — embed, cosineSim, similar — that works in Node 18+, browsers, Cloudflare Workers, Vercel Edge, Deno, and Bun with no network call.
What it is useful for
- Search-as-you-type in a browser or extension, where any network round-trip would dominate the response time.
- Privacy-sensitive and offline apps where the text must never leave the device, such as browser extensions, Obsidian plugins, or desktop tools.
- Edge runtimes (Cloudflare Workers, Vercel Edge, Deno Deploy) that embed co-located with the request handler, removing a separate inference service.
- Static sites built with Jekyll, Hugo, or Astro that ship the model inside the bundle and offer semantic search without a backend.
- IoT and small ARM devices (Raspberry Pi, gateways, kiosks) that have no GPU or NPU but still want semantic matching.
Output is a 384-dim L2-normalized Float32Array; both tiers share that shape, so the quality tier can replace the smaller one without code changes. The trade-off versus the full MiniLM-L6 teacher is modest: the README reports Spearman 0.820 (mini) and 0.844 (base), with SciFact NDCG@10 of 0.439 and 0.465.
How engineers can use it
Install one of two npm packages; the API is identical:
npm install @ternlight/base # 7.2 MB wire, ~5 ms/embed
npm install @ternlight/mini # 5.0 MB wire, ~2.5 ms/embed
Then call embed and similar directly:
import { embed, similar } from '@ternlight/base';
const hits = similar('I want my money back', [
'Refunds: how to get your money back',
'Track the status of your delivery',
'Update your billing address',
], { topK: 2 });
Bundler notes from the README:
- Vite ≥ 8.1: add the package to
optimizeDeps.exclude. Older Vite also requiresvite-plugin-wasm. - Next.js: set
experiments.asyncWebAssembly = trueinnext.config.jsand import the package from a Client Component (ornext/dynamic). - webpack 5: enable
experiments.asyncWebAssembly. - Cloudflare Workers, Vercel Edge, Deno, Bun: no extra configuration; the package routes to the right loader.
Documented limitations to plan around:
- Inputs are capped at 128 tokens (~95 words), so long documents need to be chunked first.
- It is a two-layer distilled student, not a drop-in replacement for the full
all-MiniLM-L6-v2on every domain. - The latency and size numbers were measured on an M-series Mac with Node/V8; other hardware will differ.