← Back to archive

Repo of the Day

NandhaKishorM/laya

Published: Sep 19, 2026

Open repository ↗

Contribute to NandhaKishorM/laya development by creating an account on GitHub.

Summary

Laya is a Python decision engine that classifies, scores, or answers yes/no questions about arbitrary text in a single forward pass (about 33 ms per question, 7.2 ms per question when batched on a T4), without generating any output. A built-in Router picks between an English ModernBERT-large checkpoint and a 322M multilingual model covering 100+ languages, so the same code works across locales. Decisions are typed (choice, score, noul), so there is nothing to parse or hallucinate.

What it is useful for

Laya fits well when you need structured, machine-readable decisions from text: support email triage, ticket routing, intent classification, urgency scoring, or any workflow where you currently prompt a generative model and parse its free-form answer back into JSON. It runs non-autoregressively, so it is faster and cheaper than asking an LLM for the same answer, and the typed outputs slot into LangChain, LlamaIndex, CrewAI, or an MCP server through optional extras. Multilingual routing means you do not maintain separate models per language, and the laya[serve] and laya[mcp] extras give you a local HTTP endpoint without a separate stack.

How engineers can use it

Install with python -m pip install laya (Python 3.10+). Framework wrappers are available as extras: laya[langchain], laya[llamaindex], laya[crewai]. For a server use laya[serve] or laya[mcp]; for ONNX use laya[onnx]; for a TileLang GPU fast path use laya[fast]. With uv, run uv add laya.

Define a question schema and call Router().predict:

from laya import Router

router = Router()  # downloads a checkpoint on first use
questions = {
    "department": {"type": "choice",
                   "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency":    {"type": "score",
                   "instructions": "How urgent is this?",
                   "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul",
                   "instructions": "Does the user threaten to cancel?"},
}
result = router.predict(
    "We were billed twice for March. Please refund today or we cancel.",
    questions,
)
# result["answers"]["department"]["choice"] -> "billing"
# result["answers"]["churn_risk"]["noul"]   -> probability the answer is "yes"
# result["routing"]["model"]                -> "english" or "multilingual"

A CLI preset, laya "My payment failed twice" --preset triage, answers a ready-made question set for quick checks.

The shipped checkpoints work zero-shot, but fine-tuning on your own labels is where accuracy jumps. The README reports a fine-tuned laya-typed-decisions checkpoint reaching 0.766 accuracy against 0.362 for the base English model on a 2,000-decision typed-decisions benchmark; a Kaggle 2xT4 notebook and an Apple Silicon script in notebooks/ run the full loop (build dataset, train with RLCD, calibrate temperatures, evaluate, export).

Documented limitations worth noting up front: laya-multilingual ships with a 1,024-token limit, so pass max_len=8192 for long documents; the README reports 16 to 18 of 20 correct up to about 4,000 tokens, then 8 to 17 of 20 beyond, and recommends validating on your own long documents. The Java (laya-java/) and .NET (laya-dotnet/) ports are marked advisory parity, the .NET port has no NuGet package, and a TypeScript client ships as the laya-ts npm package.