Leanpub Header

Skip to main content

Guide to Decision Models

What the "JEV-like" family is, how it is used, and how to run, deploy, and train local variants

Guide to Decision Models

Decision models answer a typed question with a calibrated probability in one forward pass: no token stream, no parsing, no values outside your schema. This guide explains TypeSafe's Jev and the open "JEV-like" family, then shows how to choose, run, serve, fine-tune and evaluate local variants on your own hardware.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
About

About

About the Book

A card transaction arrives and your system has to decide whether it looks fraudulent. Most teams ask a chat model. It generates the answer one token at a time, takes seconds, bills you for output, and can still hand back a string your code has to parse, validate and sometimes reject.

A decision model skips the generation step. It reads a state and a set of typed questions and returns a value from your schema with a calibrated probability, in a single forward pass. TypeSafe named the category System One and shipped Jev as its first model. Open variants followed: Kev, Laya, Clef and others, served locally by tools such as Ollaya and, increasingly, llama.cpp. Practitioners call the family "JEV-like".

Knowing the category exists does not tell you which checkpoint you may ship, how to serve it offline, or how to prove a local variant is good enough to replace the generative call. This book assembles the scattered launch posts, model cards and repositories into something you can run:

  • What a decision model is, what its type-safety guarantee covers, and what it does not.
  • Where System One models sit next to System Two reasoning models, and how to route between them in an agent stack.
  • Which tasks suit a bounded decision (classification, routing, scoring, safety checks), with latency and cost budgets.
  • A catalog of local variants compared by size, backbone, benchmark scores and, above all, license.
  • Running local inference, quantizing to GGUF, ONNX, Core ML and MLX, and serving behind a TypeSafe-compatible endpoint with Ollaya or kev.serve.
  • Building decision datasets, fine-tuning with LoRA and QLoRA, and measuring accuracy and calibration (ECE, Brier) with regression tests in CI.

The book is written for applied ML engineers and builders who ship models offline or on-premises and are comfortable with Python, the terminal and containers. No prior experience with decision models is assumed.

Researched and drafted with AI assistance by Ground Truth Books. Every factual claim is cited to its source (more than 200 references), and much of the performance data comes from the vendors themselves, so the book reports it as claims to test on your own workload. Commands and code follow the projects' own documentation and are cited; they were not executed for this book. Product names are used descriptively; this book is not affiliated with or endorsed by TypeSafe, Cloudflare, OpenAI, Ollama or the Ollaya project.

Author

About the Author

Ground Truth Books

Ground Truth Books publishes practical technical books on computer science, IT tools, data science, machine learning, software engineering and AI.

The name comes from machine learning, where "ground truth" means the real, verified answers you check a model against. That's how these books are made. They're researched and drafted with AI assistance, and then checked against reality. Examples are run against real captures and data in a lab, and those runs are cited like any other source, so you can see which output came straight from the tool. Every claim is cited to its source, with official documentation, specifications and source code preferred over blog posts.

Each book is built around doing the work. Chapters end with exercises, and an appendix gives worked answers you can reproduce on your own machine, using the same freely available data and captures.

Books are updated when the tools change. If you find an error, please report it: a corrected edition is free for every reader, which is one of the best things about Leanpub.

Contents

Table of Contents

Introduction

  1. Problem this book solves
  2. Who this book is for
  3. How to read this book
  4. Pinned versions and reproducibility
  5. Part I: Foundations

What a Decision Model Is

  1. Typed answers instead of text
  2. Single forward pass decisions
  3. State and typed questions
  4. Confidence and calibration
  5. The JEV name and Clef
  6. What decision models are not

System 1 and System 2

  1. Dual-process theory as architecture
  2. Reflexive versus deliberative routing
  3. Where decision models win
  4. Where deliberation is still required
  5. Part II: Using Decision Models

Task Taxonomy and Budgets

  1. Classification and intent detection
  2. Routing and tool selection
  3. Scoring and safety checks
  4. Latency and throughput budgets
  5. Cost per decision

Decision Models in Agent Stacks

  1. Decision models inside agent loops
  2. Routers and cascades
  3. Tool and skill selection
  4. Confidence as an operating policy
  5. Extraction and structured output

The Local Variant Catalog

  1. Ollaya’s listed models
  2. Typed-decisions scores
  3. Licenses and commercial use
  4. Adjacent decision backbones
  5. Multilingual checkpoints
  6. Choosing a checkpoint
  7. Part III: Running and Deploying Locally

Running Local Inference

  1. The local loop
  2. Loading a checkpoint
  3. Your first choice question
  4. Scoring a candidate answer
  5. Reading the confidence value
  6. Reproducing a quickstart
  7. Debugging bad inputs

Quantized Formats and Exports

  1. GGUF quantization tiers
  2. Converting and publishing GGUFs
  3. ONNX and runtime export
  4. Core ML and MLX targets
  5. Choosing a quant level

Serving with Ollaya

  1. Pulling and serving models
  2. Modelfiles for decision models
  3. The Ollaya MCP server
  4. ONNX Runtime versus llama.cpp
  5. kev.serve on CUDA and MLX

Serving with the TypeSafe API

  1. Typed questions over HTTP
  2. The wire API shape
  3. Using the typesafe_sdk client
  4. Pointing clients at local endpoints
  5. OpenAI’s Decisions API
  6. Gateways and drop-in compatibility
  7. Part IV: Training and Evaluating

Building Decision Datasets

  1. Intent corpora and benchmarks
  2. Synthetic decision triples
  3. Few-shot and SetFit sets
  4. Multilingual decision data
  5. Label quality and splits

Fine-Tuning with LoRA and QLoRA

  1. Parameter-efficient fine-tuning basics
  2. LoRA and QLoRA mechanics
  3. Training with Unsloth
  4. Training with Axolotl and TRL
  5. Training a decision head
  6. Merging and exporting adapters

Evaluation and Calibration

  1. Decision accuracy versus schema validity
  2. Calibration and ECE
  3. The typed-decisions benchmark
  4. The Bespoke Labs benchmark
  5. Building an eval harness
  6. Regression tests in CI

Conclusion

  1. What you can now do
  2. Choosing your fast path
  3. Where to go next

Appendix A: Command and flag reference

  1. Serving
  2. The Ollaya CLI
  3. Environment variables
  4. Training, calibration and evaluation

Appendix B: Pinned versions register

  1. What the register records
  2. Checkpoints: Kev 1.0
  3. Runtimes and serving builds
  4. Wire-format pins
  5. Benchmark pins
  6. Dataset and gold pins
  7. Training pins
  8. License pins

Appendix C: Quantization format reference

  1. Precision formats
  2. GGUF and the GGML block schemes
  3. What quantization costs a decision model
  4. Training-time formats
  5. Host formats and hardware quants
  6. Format by deployment

Glossary

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub