Kev: an open decision-model family you can train yourself

Published Sep 30, 2026

TLDR

Kev is an open source decision model that answers yes/no, menu-choice, and score questions about a short state (a ticket, a form, an agent’s context) in one request, with a probability attached to each answer. It enables software to make its small judgment calls on your own hardware: route the ticket, decide whether to escalate, score the frustration, no essay and no parsing step. The difference between Kev and regular LLMs is that a chat model writes an answer you dig out of text while Kev returns the answer itself, and the full training recipe ships with it, so you can retrain it on your own categories in about 20 minutes on one H100.

sources: - url: https://github.com/jaredpalmer/kev already-on-hn: yes (Show HN, 98 points) original-facts: - Fit math against our tracked hardware: the 4B at bf16 needs 8GB and the 9B needs 18GB, so both fit a 32GB Mac; they also clear our M4 Max 64GB, M5 Ultra 96GB, DGX Spark 128GB, and RTX 3090 24GB (6GB spare) tiers, where the repo publishes no hardware guidance. - The latency table read correctly, with the catches the repo states but buries: the Qwen3.5 generation is 3 to 6.7 times slower than the Qwen3 one on Mac (779ms vs 174ms at 4B) because no fast kernels exist yet for the DeltaNet layers, and the MLX backend is the next planned change. - The fit-checker arithmetic Tokenstead uses for every model page: total_params x bits/8, which puts Kev in the cheap tier (1.6GB at 0.8B, 8GB at 4B, 18GB at 9B) alongside Laya’s 843MB, not the 24GB-plus tier of full chat models. - A three-way comparison against our existing coverage that the repo does not attempt: Jev is closed with an API, Laya is 33ms and non-autoregressive, Kev is trainable and API-compatible with TypeSafe’s System One API, so each holds a different corner. - What the option-order playground means for calibration trust: if shuffling the answer choices changes the answer, the confidence number attached is not yet threshold-grade, and you should calibrate against your own labels before routing on it. - Where Kev lands in our decision-model suite, linking what-is-jev-plain-english, jev-typed-decisions-openjev-in-browser, and laya-multilingual-decision-model, the repo offers no such map.

What Kev is

Jared Palmer’s Kev went up as a Show HN (98 points as of today). Three sizes, 0.8B, 4B, and 9B, live on Hugging Face as jaredpalmer/kev-0.8b, kev-4b, and kev-9b, all Apache-2.0, with the frozen eval suites published as a dataset (jaredpalmer/kev-suites) so any build can be scored against the same yardstick.

Under the hood it is a LoRA adapter plus a small readout head on Qwen3.5 bases. That construction is why the training recipe matters: the weights you download are a patch on a base model, and the patch is what you retrain.

The API, and what you can do with it today

The API matches TypeSafe’s System One API (docs.typesafe.ai/api), so the TypeSafe Python SDK points at a local Kev server and works. Three question shapes ride along in one request: yes/no, choice, score. The questions share the input text and cannot read each other. TypeSafe calls its yes/no shape Noul (yes, really); plain words work fine.

The README’s example, a support ticket with three questions:

ticket: "My invoice shows $340 but my card was charged $420."
questions:
  department: billing, returns, shipping
  escalate: yes or no
  frustration: 0 to 2

Back comes one response:

department: returns 0.47, shipping 0.28, billing 0.25
escalate: 0.93
frustration: 1.44 of 2
latency_ms: 495

Your code reads it like a hash: escalate when the number clears your threshold, route on the department. The 0.93 on escalate is a clear instruction; the 0.47 on department is a soft one, and the routing logic has to tell those two cases apart.

The repo also ships a web playground that shuffles the order of the options and shows what happens to the answers. Run it against your real questions before Kev routes anything on its own. If the answer moves when the menu order changes, the confidence number attached is not yet threshold-grade: keep the option order fixed, or send low-margin answers to a human.

The latency numbers, published honestly

The repo publishes M5 numbers (bf16, 5 questions, 3 options, a state of about 230 tokens) next to the previous generation, same Mac:

model M5 latency previous generation (Qwen3 bases)
Kev-0.8B 329ms 0.6B: 123ms
Kev-4B 779ms 4B: 174ms
Kev-9B about 2s 8B: about 300ms

These are the honest numbers, published as measured. The Qwen3.5 generation is slower on Mac because no fast kernels exist yet for the DeltaNet layers; the code runs PyTorch reference implementations, and the MLX backend is the next planned change. On CUDA with flash-linear-attention, 5 questions take tens of milliseconds on an H100. Read the table as a hardware story, not a model story.

Prefix caching, from the previous generation: a repeated 772-token state answers in 242ms instead of 861ms, a 3.56x cut for the repeated-inbox case where the state stays and only the questions change.

Training your own

The recipe is the part Jev cannot offer. The decision-v7 suite: 10,000 examples from 10 public datasets, 896 generated policy examples, and 1,680 from 60 generated rule structures (12,576 total). LoRA rank 16, cross-entropy loss, 2 epochs. About 20 minutes on one H100 for the 0.8B.

Your version is your own decision suite: your question shapes, your edge cases, your 11-team department taxonomy instead of 3. You generate the examples, run the recipe, get the model, then check it against the frozen eval suites to confirm the general case did not get worse while your categories got better.

Where it fits against Jev and Laya

  Jev (TypeSafe) Laya Kev
weights closed, API only open, Apache 2.0 open, Apache 2.0
speed 70 to 500ms (published) 33ms (self-measured) 329ms to 2s on M5, tens of ms on H100
trainable no the published checkpoints are the product yes, recipe and evals public
API System One API pip package matches System One API
cost $0.042 per million input tokens free self-hosted free self-hosted

Kev holds the corner the other two do not: API-compatible with TypeSafe, and trainable by you. Jev is the managed option, Laya the fast multilingual one, Kev the one you retrain on your own categories. Full context lives in what a decision model is, in plain English, the deep Jev guide with the browser reproduction and the eight-job split, and Laya with the 33ms claims and the multilingual angle.

The fit math

The fit checker on this site runs one formula: total params x bits/8, against pooled memory. Kev at bf16:

size bf16 weights fits where
0.8B 1.6GB anything, phones included
4B 8GB 32GB Mac, M4 Max 64GB, 3090
9B 18GB 32GB Mac (tight), M4 Max 64GB, M5 Ultra 96GB, DGX Spark 128GB, 3090 24GB

The 9B is the interesting one for this audience: 18GB of weights leaves 6GB of headroom on a 24GB 3090 and runs easily on the M4 Max 64GB and M5 Ultra 96GB builds we track. Every Mac tier in our tables can run the whole family, which is not true of most model families we cover.

The honest catches

Mac is slower this generation. No fast kernels exist for the DeltaNet layers on Apple Silicon yet, so this generation runs 3 to 6.7 times slower than the previous one until the MLX backend ships. If Mac latency is a hard requirement, the Qwen3-based checkpoints are faster today.

The probabilities drift. bf16 versus fp32, across 24 records, moves the numbers by at most 0.017, and the top answer never changes. Route on the top answer and the gap to second place; do not build logic that cares about the fourth decimal.

Why the category keeps compounding

Think of it like a coffee order at a busy counter: the menu is printed, you say the order, you pay, you move on. Kev is the printed menu of models, the question shapes fixed before you walk up. Jev is the same menu at their counter, and you pay per order. Kev hands you the print shop too: change the menu, print it, the counter is yours.

Three weeks ago this layer was one closed API. Then a browser reproduction (OpenJev), then a 33ms multilingual encoder (Laya), now a trainable open family with the recipe published (Kev). Each step moved something from the vendor’s side of the line to yours: the runtime, the languages, now the training. The practical state for a local builder: the answers cost nothing, they run on hardware you already own at these sizes, and the model behind them is now something you retrain instead of something you wait on.

Discussion

Be the first to comment

Start a discussion

Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.