A card transaction arrives and your system has to decide whether it looks fraudulent. Most teams ask a chat model. It generates the answer one token at a time, takes seconds, bills you for output, and can still hand back a string your code has to parse, validate and sometimes reject.
A decision model skips the generation step. It reads a state and a set of typed questions and returns a value from your schema with a calibrated probability, in a single forward pass. TypeSafe named the category System One and shipped Jev as its first model. Open variants followed: Kev, Laya, Clef and others, served locally by tools such as Ollaya and, increasingly, llama.cpp. Practitioners call the family "JEV-like".
Knowing the category exists does not tell you which checkpoint you may ship, how to serve it offline, or how to prove a local variant is good enough to replace the generative call. This book assembles the scattered launch posts, model cards and repositories into something you can run:
- What a decision model is, what its type-safety guarantee covers, and what it does not.
- Where System One models sit next to System Two reasoning models, and how to route between them in an agent stack.
- Which tasks suit a bounded decision (classification, routing, scoring, safety checks), with latency and cost budgets.
- A catalog of local variants compared by size, backbone, benchmark scores and, above all, license.
- Running local inference, quantizing to GGUF, ONNX, Core ML and MLX, and serving behind a TypeSafe-compatible endpoint with Ollaya or kev.serve.
- Building decision datasets, fine-tuning with LoRA and QLoRA, and measuring accuracy and calibration (ECE, Brier) with regression tests in CI.
The book is written for applied ML engineers and builders who ship models offline or on-premises and are comfortable with Python, the terminal and containers. No prior experience with decision models is assumed.
Researched and drafted with AI assistance by Ground Truth Books. Every factual claim is cited to its source (more than 200 references), and much of the performance data comes from the vendors themselves, so the book reports it as claims to test on your own workload. Commands and code follow the projects' own documentation and are cited; they were not executed for this book. Product names are used descriptively; this book is not affiliated with or endorsed by TypeSafe, Cloudflare, OpenAI, Ollama or the Ollaya project.