← Back to archive

Repo of the Day

soycaporal/ternlight

Published: Aug 30, 2026

Open repository ↗

Contribute to soycaporal/ternlight development by creating an account on GitHub.

Summary

ternlight packages a small sentence-embedding model, BERT-style tokenizer, and inference engine into a single 5–7 MB WebAssembly file that runs on CPU. The weights are ternary (-1, 0, +1) and the Rust engine is compiled to WASM SIMD, so embeddings reduce to adds and subtracts rather than floating-point multiplies. It exposes a tiny JavaScript API — embed, cosineSim, similar — that works in Node 18+, browsers, Cloudflare Workers, Vercel Edge, Deno, and Bun with no network call.

What it is useful for

  • Search-as-you-type in a browser or extension, where any network round-trip would dominate the response time.
  • Privacy-sensitive and offline apps where the text must never leave the device, such as browser extensions, Obsidian plugins, or desktop tools.
  • Edge runtimes (Cloudflare Workers, Vercel Edge, Deno Deploy) that embed co-located with the request handler, removing a separate inference service.
  • Static sites built with Jekyll, Hugo, or Astro that ship the model inside the bundle and offer semantic search without a backend.
  • IoT and small ARM devices (Raspberry Pi, gateways, kiosks) that have no GPU or NPU but still want semantic matching.

Output is a 384-dim L2-normalized Float32Array; both tiers share that shape, so the quality tier can replace the smaller one without code changes. The trade-off versus the full MiniLM-L6 teacher is modest: the README reports Spearman 0.820 (mini) and 0.844 (base), with SciFact NDCG@10 of 0.439 and 0.465.

How engineers can use it

Install one of two npm packages; the API is identical:

npm install @ternlight/base    # 7.2 MB wire, ~5 ms/embed
npm install @ternlight/mini    # 5.0 MB wire, ~2.5 ms/embed

Then call embed and similar directly:

import { embed, similar } from '@ternlight/base';

const hits = similar('I want my money back', [
  'Refunds: how to get your money back',
  'Track the status of your delivery',
  'Update your billing address',
], { topK: 2 });

Bundler notes from the README:

  • Vite ≥ 8.1: add the package to optimizeDeps.exclude. Older Vite also requires vite-plugin-wasm.
  • Next.js: set experiments.asyncWebAssembly = true in next.config.js and import the package from a Client Component (or next/dynamic).
  • webpack 5: enable experiments.asyncWebAssembly.
  • Cloudflare Workers, Vercel Edge, Deno, Bun: no extra configuration; the package routes to the right loader.

Documented limitations to plan around:

  • Inputs are capped at 128 tokens (~95 words), so long documents need to be chunked first.
  • It is a two-layer distilled student, not a drop-in replacement for the full all-MiniLM-L6-v2 on every domain.
  • The latency and size numbers were measured on an M-series Mac with Node/V8; other hardware will differ.