Hands-On LLM Serving and Optimization : Hosting LLMs at Scale
Book Details
Format
Paperback / Softback
ISBN-10
834162149Y
ISBN-13
9798341621497
Publisher
O'Reilly Media
Imprint
O'Reilly Media
Country of Manufacture
GB
Country of Publication
GB
Publication Date
May 26th, 2026
Print length
300 Pages
Weight
656 grams
Dimensions
17.60 x 23.40 x 2.20 cms
Product Classification:
Natural language & machine translationNatural language and machine translation
Ksh 11,500.00
Werezi Extended Catalogue
Delivery in 14 days
10 copies in stock
Delivery Location
Delivery fee: Select location
Delivery in 14 days
Secure
Quality
Fast
As the demand for real-time AI applications grows, along comes this comprehensive guide to the complexities of deploying and optimizing LLMs at scale. The authors take a real-world approach backed by practical examples and code, and assemble essential strategies for designing infrastructures that are equal to the demands of modern AI applications.
Large language models (LLMs) are rapidly becoming the backbone of AI-driven applications. Without proper optimization, however, LLMs can be expensive to run, slow to serve, and prone to performance bottlenecks. As the demand for real-time AI applications grows, along comes Hands-On Serving and Optimizing LLM Models, a comprehensive guide to the complexities of deploying and optimizing LLMs at scale. In this hands-on book, authors Chi Wang and Peiheng Hu take a real-world approach backed by practical examples and code, and assemble essential strategies for designing robust infrastructures that are equal to the demands of modern AI applications. Whether you're building high-performance AI systems or looking to enhance your knowledge of LLM optimization, this indispensable book will serve as a pillar of your success. Learn the key principles for designing a model-serving system tailored to popular business scenariosUnderstand the common challenges of hosting LLMs at scale while minimizing costsPick up practical techniques for optimizing LLM serving performanceBuild a model-serving system that meets specific business requirementsImprove LLM serving throughput and reduce latencyHost LLMs in a cost-effective manner, balancing performance and resource efficiency
Get Hands-On LLM Serving and Optimization by at the best price and quality guaranteed only at Werezi Africa's largest book ecommerce store. The book was published by O'Reilly Media and it has pages.