A practical guide to evaluating small and local language models in the real world. Learn how to benchmark reliably, compare models across hardware and runtimes, optimize performance and avoid misleading results. With practical code and rigorous methods, you can build your own evaluation setup and make confident deployment decisions.