Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Roland Huss, Daniele Zonca

Rating No ratings yet

Generative AI is revolutionizing industries, and Kubernetes has fast become the backbone for deploying and managing these resource-intensive workloads. This book serves as a practical, hands-on guide for MLOps engineers, software developers, Kubernetes administrators, and AI professionals ready to combine AI innovation with the power of cloud native infrastructure. Authors Roland Huß and Daniele Zonca provide a clear road map for training, fine-tuning, deploying, and scaling GenAI models on Kubernetes, addressing challenges like resource optimization, automation, and security along the way. With actionable insights with real-world examples, readers will learn to tackle the opportunities and complexities of managing GenAI applications in production environments. Whether you're experimenting with large-scale language models or facing the nuances of AI deployment at scale, you'll uncover expertise you need to operationalize this exciting technology effectively. Learn how to deploy LLMs more efficiently with optimized inference runtimes Get hands-on with GPU scheduling, including hardware detection and multinode scaling Monitor and understand LLM-specific metrics like Time to First Token and token throughput Know when to fine-tune a model or when retrieval augmentation is the better choice Discover how to evaluate models with standardized benchmarks before committing GPU resources Learn to run agentic applications with secure tool integration, identity management, and persistent state

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, hands-on guide for MLOps engineers, Kubernetes administrators, and AI professionals who want to deploy, scale, and operate large language models (LLMs) on Kubernetes, covering everything from GPU scheduling to model evaluation and agentic applications. 【Book Arc】 - **Opening (~0%–9%)**: Introduces the core premise—Kubernetes as the ideal platform for GenAI workloads—and provides a brief history of generative AI, from early concepts like the Turing Test to the deep learning revolution. Sets up the book's scope: training, fine-tuning, deploying, and scaling LLMs. - **Early (~9%–25%)**: Explains the fundamentals of LLMs, including tokenization, prompt engineering, and the Transformer architecture. Covers how to run models locally with Hugging Face Transformers, highlighting hardware requirements (e.g., GPU memory for 7B vs. 70B models) and the role of runtimes like TGI. - **Early (~25%–38%)**: Moves into Kubernetes-specific deployment, contrasting manual approaches (with taints, tolerations, PVCs) against higher-level abstractions like model server controllers (e.g., KServe's LLMInferenceService). Discusses distributed inference patterns for large models (70B+) and the trade-offs of single-node vs. multinode serving. - **Middle (~38%–53%)**: Focuses on model data management—storage formats (GGUF, Safetensors, ONNX), model registries (like MLflow), and how to organize, version, and serve model artifacts within a cluster. Emphasizes the separation of metadata from actual model weights for flexibility. - **Late (~53%–end)**: Covers production concerns: GPU resource scheduling (including NVIDIA GPU Operator and dynamic allocation), autoscaling, LLM-specific metrics (Time to First Token, token throughput), model compression, evaluation benchmarks, and advanced topics like agentic applications with secure tool integration and persistent state. 【Key Takeaways】 - **Kubernetes is the right platform for GenAI** (Early): Unlike specialized tools like Ray or Spark, Kubernetes can manage AI models alongside traditional applications, databases, and microservices, simplifying operations and enabling complex multi-workload compositions. - **Subword tokenization is essential for LLMs** (Early): Breaking text into subword units (e.g., "regularization" → "regular" + "ization") eliminates the unknown-word problem while keeping vocabulary sizes manageable, allowing models to handle rare terms and technical jargon. - **GPU memory scales with model size** (Early): A 7B-parameter model requires ~15 GB of GPU memory, while a 70B model needs ~140 GB. This drives decisions about hardware, quantization, and whether to use single-node or distributed serving. - **Model server controllers abstract Kubernetes complexity** (Early): Manual deployment of LLMs involves managing taints, tolerations, storage, secrets, and model-specific parameters—complexity that multiplies with each model. Controllers like KServe's LLMInferenceService hide this behind higher-level APIs focused on model deployment. - **Model formats are still evolving** (Middle): GGUF and Safetensors are currently the most practical choices for balancing performance, compatibility, and flexibility, but true standardization (like OCI for containers) is still far off. Safetensors' multifile structure works well with OCI artifacts for efficient caching and parallel downloads. - **Model registries separate metadata from weights** (Middle): Registries like MLflow manage model versions, governance, and metadata while referencing external object stores (e.g., S3) for actual weights. This separation enables flexibility and keeps metadata accessible within the cluster. - **LLM-specific metrics matter for production** (Late): Monitoring Time to First Token and token throughput is critical for optimizing user experience and resource utilization, requiring rethinking traditional autoscaling and traffic management for long-running GPU workloads. 【Reading Tips】 - **Skim the history and fundamentals** (Early): The opening chapters on GenAI history and Transformer architecture are useful context but not the core value. Focus on the practical deployment examples and code snippets. - **Deep-read the GPU scheduling and model server controller sections** (Early–Middle): These are the heart of the book—understanding taints/tolerations, dynamic resource allocation, and KServe's LLMInferenceService will save you hours of trial-and-error in production. - **Pay attention to model format comparisons** (Middle): The discussion of GGUF vs. Safetensors vs. ONNX is practical and decision-oriented. Use it to choose the right format for your runtime and storage strategy. - **Treat the production chapter as a checklist** (Late): Autoscaling, vLLM tuning, and evaluation benchmarks are actionable items you'll want to revisit when deploying your own models. Consider bookmarking the vLLM runtime parameters section. - **Note the gaps**: The excerpts don't cover agentic applications in depth, nor do they detail specific security implementations. If those are your focus, you may need supplementary resources. 【Coverage Limits】 This guide synthesizes the first ~53% of the book (through model registries). Later sections on production tuning, autoscaling, and agentic applications are only partially covered based on available excerpts.
Excerpt 1
48 MLflow Model Registry 49 Kubeflow Model Registry 54 OCI Registry 57 Accessing Model Data in Kubernetes 59 Shared Storage with PersistentVolumes 61 OCI Ima...
View in text
Excerpt 2
stem prompt that guides the LLM behavior. The system prompt is included in the full request and defines the scenario that the model should use to handle the...
View in text
Excerpt 3
rmers, peft, or diffusers, are incubated in this community. TGI now supports multiple inference backends, allowing you to choose the most appropriate backend...
View in text
Excerpt 4
rganizations deploy most model registries as local services within a cluster. Organizations don’t expose these registries outside the cluster. The registries...
View in text
Excerpt 5
odelcar containers starting quickly, startup is slower when the modelcar image still needs to be pulled from an OCI Registry. This can be mitigated by using...
View in text
Excerpt 6
ngle-Node Versus Multinode Inference” on page 110. In prac‐ tice, the maximum tensor parallel degree is often the number of GPUs in one server (e.g., four-wa...
View in text
Excerpt 7
t it exposes an OpenAI-compatible API. vLLM benchmark suite Performance is a critical aspect for inference engines like vLLM, which provides publicly availab...
View in text
Excerpt 8
tance to another is very fast, in the range of milliseconds. The idea is quite natural, but the implementation is very complex, and two new projects have bee...
View in text
Tags
AI categories
Cloud NativeAIBackend
ISBN: 1098171926
Publisher: O'Reilly Media
Publish Year: 2026
Language: English
Pages: 407
File Format: PDF
File Size: 8.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…