Applied AI Engineering
Retrieval, Fine-Tuning, Evaluation, and Deployment Patterns for Real-World LLM Systems
-
- £12.99
-
- £12.99
Publisher Description
Stop building demos. Start building systems that survive real users.
The gap between a working prototype and a production AI system is the hardest part of engineering with large language models, and it is the part most resources gloss over. This book fills that gap with battle-tested patterns for retrieval, fine-tuning, evaluation, and deployment that have survived contact with real users.
Inside, you will learn:
When to use retrieval versus fine-tuning, and how to do both
How to build production RAG systems that actually work
Advanced retrieval patterns including hybrid search, reranking, and iterative retrieval
Fine-tuning strategy and data preparation for real results
Parameter-efficient fine-tuning with LoRA and QLoRA
Evaluation as an engineering practice, including LLM-as-judge and critique shadowing
Deployment patterns including caching, batching, and model routing
Cost optimization and intelligent model routing strategies
Guardrails and safety architecture for production systems
Monitoring and observability for LLM systems
This is not a theoretical book. Every recommendation comes from direct experience building production LLM systems across multiple organizations. When the book tells you something works, it is because that approach has survived contact with real users. When it tells you something does not work, it is because the author has seen the failure firsthand.
Whether you are a software engineer building AI features for the first time, an engineering leader making strategic decisions, or an experienced ML practitioner adapting to LLMs, this book gives you the decision frameworks you need.
The landscape of AI tools changes constantly, but the patterns in this book endure. Start with retrieval. Measure your errors. Use fine-tuning to close the gaps. Deploy with confidence. That is the pattern that never goes out of style.