Generative AI is revolutionizing industries, and Kubernetes has fast become the backbone for deploying and managing these resource-intensive workloads. This book serves as a practical, hands-on guide for MLOps engineers, software developers, Kubernetes administrators, and AI professionals ready to combine AI innovation with the power of cloud native infrastructure. Authors Roland Huß and Daniele Zonca provide a clear road map for training, fine-tuning, deploying, and scaling GenAI models on Kubernetes, addressing challenges like resource optimization, automation, and security along the way.
With actionable insights with real-world examples, readers will learn to tackle the opportunities and complexities of managing GenAI applications in production environments. Whether you're experimenting with large-scale language models or facing the nuances of AI deployment at scale, you'll uncover expertise you need to operationalize this exciting technology effectively.
Learn how to deploy LLMs more efficiently with optimized inference runtimes
Get hands-on with GPU scheduling, including hardware detection and multinode scaling
Monitor and understand LLM-specific metrics like Time to First Token and token throughput
Know when to fine-tune a model or when retrieval augmentation is the better choice
Discover how to evaluate models with standardized benchmarks before committing GPU resources
Learn to run agentic applications with secure tool integration, identity management, and persistent state
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical, hands-on guide for MLOps engineers, Kubernetes administrators, and AI professionals who want to deploy, scale, and operate large language models (LLMs) on Kubernetes, covering everything from GPU scheduling to model evaluation and agentic applications.
【Book Arc】
- **Opening (~0%–9%)**: Introduces the core premise—Kubernetes as the ideal platform for GenAI workloads—and provides a brief history of generative AI, from early concepts like the Turing Test to the deep learning revolution. Sets up the book's scope: training, fine-tuning, deploying, and scaling LLMs.
- **Early (~9%–25%)**: Explains the fundamentals of LLMs, including tokenization, prompt engineering, and the Transformer architecture. Covers how to run models locally with Hugging Face Transformers, highlighting hardware requirements (e.g., GPU memory for 7B vs. 70B models) and the role of runtimes like TGI.
- **Early (~25%–38%)**: Moves into Kubernetes-specific deployment, contrasting manual approaches (with taints, tolerations, PVCs) against higher-level abstractions like model server controllers (e.g., KServe's LLMInferenceService). Discusses distributed inference patterns for large models (70B+) and the trade-offs of single-node vs. multinode serving.
- **Middle (~38%–53%)**: Focuses on model data management—storage formats (GGUF, Safetensors, ONNX), model registries (like MLflow), and how to organize, version, and serve model artifacts within a cluster. Emphasizes the separation of metadata from actual model weights for flexibility.
- **Late (~53%–end)**: Covers production concerns: GPU resource scheduling (including NVIDIA GPU Operator and dynamic allocation), autoscaling, LLM-specific metrics (Time to First Token, token throughput), model compression, evaluation benchmarks, and advanced topics like agentic applications with secure tool integration and persistent state.
【Key Takeaways】
- **Kubernetes is the right platform for GenAI** (Early): Unlike specialized tools like Ray or Spark, Kubernetes can manage AI models alongside traditional applications, databases, and microservices, simplifying operations and enabling complex multi-workload compositions.
- **Subword tokenization is essential for LLMs** (Early): Breaking text into subword units (e.g., "regularization" → "regular" + "ization") eliminates the unknown-word problem while keeping vocabulary sizes manageable, allowing models to handle rare terms and technical jargon.
- **GPU memory scales with model size** (Early): A 7B-parameter model requires ~15 GB of GPU memory, while a 70B model needs ~140 GB. This drives decisions about hardware, quantization, and whether to use single-node or distributed serving.
- **Model server controllers abstract Kubernetes complexity** (Early): Manual deployment of LLMs involves managing taints, tolerations, storage, secrets, and model-specific parameters—complexity that multiplies with each model. Controllers like KServe's LLMInferenceService hide this behind higher-level APIs focused on model deployment.
- **Model formats are still evolving** (Middle): GGUF and Safetensors are currently the most practical choices for balancing performance, compatibility, and flexibility, but true standardization (like OCI for containers) is still far off. Safetensors' multifile structure works well with OCI artifacts for efficient caching and parallel downloads.
- **Model registries separate metadata from weights** (Middle): Registries like MLflow manage model versions, governance, and metadata while referencing external object stores (e.g., S3) for actual weights. This separation enables flexibility and keeps metadata accessible within the cluster.
- **LLM-specific metrics matter for production** (Late): Monitoring Time to First Token and token throughput is critical for optimizing user experience and resource utilization, requiring rethinking traditional autoscaling and traffic management for long-running GPU workloads.
【Reading Tips】
- **Skim the history and fundamentals** (Early): The opening chapters on GenAI history and Transformer architecture are useful context but not the core value. Focus on the practical deployment examples and code snippets.
- **Deep-read the GPU scheduling and model server controller sections** (Early–Middle): These are the heart of the book—understanding taints/tolerations, dynamic resource allocation, and KServe's LLMInferenceService will save you hours of trial-and-error in production.
- **Pay attention to model format comparisons** (Middle): The discussion of GGUF vs. Safetensors vs. ONNX is practical and decision-oriented. Use it to choose the right format for your runtime and storage strategy.
- **Treat the production chapter as a checklist** (Late): Autoscaling, vLLM tuning, and evaluation benchmarks are actionable items you'll want to revisit when deploying your own models. Consider bookmarking the vLLM runtime parameters section.
- **Note the gaps**: The excerpts don't cover agentic applications in depth, nor do they detail specific security implementations. If those are your focus, you may need supplementary resources.
【Coverage Limits】
This guide synthesizes the first ~53% of the book (through model registries). Later sections on production tuning, autoscaling, and agentic applications are only partially covered based on available excerpts.
Excerpt 1
48 MLflow Model Registry 49 Kubeflow Model Registry 54 OCI Registry 57 Accessing Model Data in Kubernetes 59 Shared Storage with PersistentVolumes 61 OCI Ima...
stem prompt that guides the LLM behavior. The system prompt is included in the full request and defines the scenario that the model should use to handle the...
rmers, peft, or diffusers, are incubated in this community. TGI now supports multiple inference backends, allowing you to choose the most appropriate backend...
rganizations deploy most model registries as local services within a cluster. Organizations don’t expose these registries outside the cluster. The registries...
odelcar containers starting quickly, startup is slower when the modelcar image still needs to be pulled from an OCI Registry. This can be mitigated by using...
ngle-Node Versus Multinode Inference” on page 110. In prac‐ tice, the maximum tensor parallel degree is often the number of GPUs in one server (e.g., four-wa...
t it exposes an OpenAI-compatible API. vLLM benchmark suite Performance is a critical aspect for inference engines like vLLM, which provides publicly availab...
tance to another is very fast, in the range of milliseconds. The idea is quite natural, but the implementation is very complex, and two new projects have bee...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Generative AI on Kubernetes Operationalizing Large Language Models (Roland Huss, Daniele Zonca)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Generative AI on Kubernetes Operationalizing Large Language Models (Roland Huss, Daniele Zonca)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment