A practical, operations-focused handbook for engineers working with HPC and AI infrastructure. Learn how to troubleshoot Linux, InfiniBand, NVIDIA GPUs, Slurm, UFM, NCCL, storage, and cluster infrastructure using a structured, evidence-driven approach.