scieee AI-readable full text Open interactive document viewer

Gen AI

Upadhyayula, Raghavender Surya

Full text

MANNING Amit Bahree Foreword by Eric Boyd Sponsored by 2EPILOGUE Input text (prompt) Token Embedding Encoder Decoder ………………….. ………………….. ………………….. Generated text (completion) Numerical representation Needed for scenarios such as “Bring your own data,” search, etc. LLM ………………….. ………………….. ………………….. ………………….. ………………….. …………………. LLM Input token vector Vector representation of next output token the mat pad … … … Highest probability Second highest probability Less likely Next word ……… ……… The dog sat on ……… ……… ……… ……… ……… ……… ……… ……… Conceptual architecture of an LLM LLM – Next token predictor Generative AI in Action AMIT BAHREE FOREWORD BY ERIC BOYD MANNING SHELTER ISLAND For online information and ordering of this and other Manning books, please visit www.manning.com. The publisher offers discounts on this book when ordered in quantity. For more information, please contact Special Sales Department Manning Publications Co. 20 Baldwin Road PO Box 761 Shelter Island, NY 11964 Email: [email protected]m ©2024 by Manning Publications Co. All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted, in any form or by means electronic, mechanical, photocopying, or otherwise, without prior written permission of the publisher. Many of the designations used by manufacturers and sellers to distinguish their products are claimed as trademarks. Where those designations appear in the book, and Manning Publications was aware of a trademark claim, the designations have been printed in initial caps or all caps. Recognizing the importance of preserving what has been written, it is Manning’s policy to have the books we publish printed on acid-free paper, and we exert our best efforts to that end. Recognizing also our responsibility to conserve the resources of our planet, Manning books are printed on paper that is at least 15 percent recycled and processed without the use of elemental chlorine. The author and publisher have made every effort to ensure that the information in this book was correct at press time. The author and publisher do not assume and hereby disclaim any liability to any party for any loss, damage, or disruption caused by errors or omissions, whether such errors or omissions result from negligence, accident, or any other cause, or from any usage of the information herein. Manning Publications Co. Development editor: Rebecca Johnson 20 Baldwin Road Technical editor: Wee Hyong Tok PO Box 761 Review editor: Radmila Ercegovac Shelter Island, NY 11964 Production editor: Kathy Rossland Copy editor: Lana Todorovic-Arndt Proofreader: Melody Dolab Technical proofreader: John Aziz Typesetter and cover designer: Marija Tudor ISBN 9781633435339 Printed in the United States of America To my family, who patiently listened to my tech rambles, although they were no help in writing this book and will never read it, and to you, dear reader, who boldly chose to engage with these ideas— may your neurons spark joy and your circuits never short. Together, let’s build a future where AI is more brains than brawn. iv brief contents PART 1FOUNDATIONS OF GENERATIVE AI ................................... 1 1■Introduction to generative AI 3 2■Introduction to large language models 26 3■Working through an API: Generating text 57 4■From pixels to pictures: Generating images 96 5■What else can AI generate? 127 PART 2ADVANCED TECHNIQUES AND APPLICATIONS 153 6■Guide to prompt engineering 155 7■Retrieval-augmented generation: The secret weapon 183 8■Chatting with your data 213 9■Tailoring models with model adaptation and fine-tuning 242 PART 3DEPLOYMENT AND ETHICAL CONSIDERATIONS 281 10 ■Application architecture for generative AI apps 283 11 ■Scaling up: Best practices for production deployment 321 12 ■Evaluations and benchmarks 357 13 ■Guide to ethical GenAI: Principles, practices, and pitfalls 384 v contents foreword xii preface xiv acknowledgments xvi about this book xviii about the author xxiii about the cover illustration xxiv PART 1 FOUNDATIONS OF GENERATIVE AI .................... 1 1 Introduction to generative AI 3 1.1 What is this book about? 5 1.2 What is generative AI? 6 1.3 What can we generate? 9 Entities extraction 9 ■Generating text 10 ■Generating images 12 ■Generating code 12 ■Ability to solve logic problems 14 ■Generating music 15 ■Generating videos 17 1.4 Enterprise use cases 17 1.5 When not to use generative AI 19 1.6 How is generative AI different from traditional AI? 19 1.7 What approach should enterprises take? 21 1.8 Architecture considerations 23 1.9 So your enterprise wants to use generative AI. Now what? 24 CONTENTSvi 2 Introduction to large language models 26 2.1 Overview of foundational models 27 2.2 Overview of LLMs 29 2.3 Transformer architecture 30 2.4 Training cutoff 31 2.5 Types of LLMs 31 2.6 Small language models 33 2.7 Open source vs. commercial LLMs 35 Commercial LLMs 36 ■Open source LLMs 36 2.8 Key concepts of LLMs 38 Prompts 39 ■Tokens 40 ■Counting tokens 42 Embeddings 45 ■Model configuration 47 ■Context window 50 ■Prompt engineering 51 ■Model adaptation 52 Emergent behavior 52 3 Working through an API: Generating text 57 3.1 Model categories 58 Dependencies 60 ■Listing models 62 3.2 Completion API 64 Expanding completions 67 ■Azure content safety filter 68 Multiple completions 69 ■Controlling randomness 71 Controlling randomness using top_p 74 3.3 Advanced completion API options 75 Streaming completions 75 ■Influencing token probabilities: logit_bias 77 ■Presence and frequency penalties 80 Log probabilities 82 3.4 Chat completion API 84 System role 86 ■Finish reason 88 ■Chat completion API for nonchat scenarios 88 ■Managing conversation 89 Best practices for managing tokens 92 ■Additional LLM providers 93 4 From pixels to pictures: Generating images 96 4.1 Vision models 97 Variational autoencoders 100 ■Generative adversarial networks 101 ■Vision transformer models 102 Diffusion models 104 ■Multimodal models 106 4.2 Image generation with Stable Diffusion 109 Dependencies 109 ■Generating an image 111 CONTENTS vii 4.3 Image generation with other providers 114 OpenAI DALLE 3 114 ■Bing image creator 114 Adobe Firefly 115 4.4 Editing and enhancing images using Stable Diffusion 116 Generating using image-to-image API 119 ■Using the masking API 121 ■Resize using the upscale API 124 ■Image generation tips 125 5 What else can AI generate? 127 5.1 Code generation 128 Can I trust the code? 130 ■GitHub Copilot 132 How Copilot works 135 5.2 Additional code-related tasks 136 Code explanation 136 ■Generate tests 138 ■Code referencing 139 ■Code refactoring 140 5.3 Other code generation tools 140 Amazon CodeWhisperer 141 ■Code Llama 142 Tabnine 143 ■Check yourself 145 ■Best practices for code generation 145 5.4 Video generation 146 5.5 Audio and music generation 149 PART 2ADVANCED TECHNIQUES AND APPLICATIONS 153 6 Guide to prompt engineering 155 6.1 What is prompt engineering? 156 Why do we need prompt engineering? 156 6.2 The basics of prompt engineering 158 6.3 In-context learning and prompting 161 6.4 Prompt engineering techniques 163 System message 163 ■Zero-shot, few-shot, and many-shot learning 166 ■Use clear syntax 168 ■Making in-context learning work 169 ■Reasoning: Chain of Thought 170 Self-consistency sampling 173 6.5 Image prompting 175 6.6 Prompt injection 176 6.7 Prompt engineering challenges 179 xiv preface With nearly 30 years of experience as a developer and applied researcher, I have been involved in fundamental technology shifts from the early days. Generative artificial intelligence (AI) is one of those areas where the hype and the fear of missing out reach stratospheric levels! Organizations are trying to understand this new technology and how to implement it. Some of this means trying to gain an edge; in other cases, it is responding to the market and the pressure from the board and CEOs to join the trend. At Microsoft, I have the privilege of being part of the Azure AI platform engineering team, helping develop some of our advanced AI technologies, such as Azure OpenAI, and Azure AI Services, including speech, vision, and small language models (e.g., the new Phi family of models). Part of my role has been collaborating with many Fortune 500 companies that are our clients. These companies are scattered around the world, representing different industry domains, with many of them being leaders in their fields. My experience with GenAI across various domains and applications, particularly in collaboration with Fortune 500 companies, has revealed that there is a gap between the hype and the reality of generative AI. I’ve noticed that many users and customers are confused or intimidated by the complexity and challenges of this field. In response, I set out to write a book to bridge this gap, providing a practical and accessible guide to generative AI. This guide empowers anyone, regardless of background, to learn and apply generative AI effectively. The technology industry is known for its rapid pace, but the field of GenAI is growing even faster, and we see changes in weeks rather than months and years. While I was writing this book, the technology advanced, and I have had to update many of the new areas in the book several times. However, the basics of GenAI and large language PREFACE xv models (LLM) remain novel and crucial to grasp. These are the building blocks on which new areas are being developed. Understanding these fundamentals is not just a goal of the book but a necessity in this rapidly evolving field. This book focuses on generative AI aspects, especially LLMs, which are often the most common use cases. I expect newer models with additional multimodal capabilities that combine vision, speech, and video will grow in the future. Here, we’ll mainly use OpenAI and Azure OpenAI, but I also show other providers’ examples. Most LLM providers are similar to OpenAI, so the book is beneficial even if you use a different provider. I also used Python for the examples, as it is easy and common in AI. In addition, there are SDKs for most languages and REST APIs that you can call in any language. Welcome to Generative AI in Action, a book aiming to demystify the generative AI field and help you apply it to your projects. I am excited to share some insights from my learning and assist you on your path. xvi acknowledgments First and foremost, I want to thank my parents for letting me disappear into the “computer room” to tinker with those amazing machines and for buying me my first computer. I also thank my wife, Meenakshi, for putting up with me, especially when I conveniently ignored most other things and worked through the graveyard shift after long days to write the book and code. To my daughter Maya, I thank you for never doubting my literal and coding abilities (even if it came with a teenager’s eye roll). This book would not be complete without my dog, Champ, who, as you will see, is a recurring theme. And finally, I thank my dear friend Somya for showing us what true courage looks like and reminding us that most of life’s dramas are just things we get ourselves worked up over. I thank Eric Boyd for writing the foreword and for his time and collaboration on this project. Working under his guidance on the Azure AI team has been an exhilarating experience. Pushing the limits of technology and rekindling that childlike excitement in all of us—it reminds me why I fell in love with computers and programming in the first place. A special thanks goes to Wee Hyong Tok, the technical editor of this book, for his incredible time spent assisting, directing, challenging, and verifying everything. Your efforts have been invaluable in my learning and in improving this book! Wee Hyong is a partner director of product at Microsoft. He has a PhD in computer science from the National University of Singapore and is a recognized expert on data and AI. He has also authored over 10 books on AI. To all the reviewers—Amit Basnak, Andres Sacco, Arun Kandregula, Bruno Ricardo Santos, Dan Sheikh, Erim Ertürk, Gregory V, Hariskumar Panakkal, Ike Okonkwo, James Coates, Julien Pohie, Lokesh Kumar, Louis Luangkesorn, Luiz Davi, Manish Jain, Matteo Battista, Maxim Volgin, Nathan B. Crocker, Pradeep Bhattiprolu, ACKNOWLEDGMENTS xvii Radhakrishna MV, Raj Kumar, Rambabu Posa, Roy Wilsker, Rui Liu, Sanjeev Jaiswal, Scott Ling, Simon Verhoeven, Sumit Pal, Sushil Singh, Swaminathan Subramanian, Swapneelkumar Deshpande, Victor Durán, and Weronika Burman—your suggestions helped make this a better book. Finally, I would like to thank the team at Manning. I have immense empathy and gratitude for my development editor, Rebecca Johnson, and acquisitions editor, Mike Stephens. Rebecca especially deserves a medal for making sense of my initial drafts and turning gibberish into coherent content. Thank you all for your patience and dedication! xviii about this book Generative AI in Action is designed to equip enterprise professionals and enthusiasts with the knowledge and skills to effectively use generative AI technologies. This book provides a comprehensive understanding of generative AI, covering its fundamental principles, practical applications, and the challenges associated with implementing it in real-world scenarios. The book teaches you how to create and use generative models for tasks and use cases. It focuses on this technology’s practical and hands-on aspects and how it works. It does not dive deep into the science, but it references the papers and scientific breakthroughs that have helped develop some of the technology—you can see these at the end of the book. This book is designed to provide a comprehensive understanding of generative AI and its potential within an enterprise context. It explores foundational models, large language models, and related algorithms and architectures, offering readers a thorough grasp of these advanced technologies. Practical insights and examples are provided to help develop and deploy generative AI models, ensuring that readers can apply these concepts in real-world scenarios. Advanced topics such as prompt engineering, retrieval-augmented generation, and model adaptation are discussed in detail, giving readers an in-depth understanding of these cutting-edge techniques. The book also highlights best practices for integrating generative AI into existing systems and workflows, ensuring a smooth and efficient implementation. Furthermore, it addresses the ethical considerations, governance, and safety measures necessary for responsible AI deployment, guiding readers on how to responsibly navigate the complexities of this rapidly evolving field. ABOUT THIS BOOK xix Who should read this book Generative AI in Action is designed for a diverse audience. It is ideal for developers and software architects looking to integrate generative AI into their projects and data scientists who want to enhance their understanding of generative AI technologies and applications. Business and technical decision-makers will find it valuable for grasping the strategic implications of generative AI for their organizations. Power users across various enterprise sectors can explore generative AI’s practical applications and benefits. Additionally, educators and students in AI-related fields will gain comprehensive knowledge of the latest advancements in generative AI. This book primarily targets developers, data scientists, and technology decisionmakers with some programming background who want to explore the fascinating and powerful world of generative AI. One doesn’t need to be an expert in machine learning, deep learning, or generative AI or have a PhD in mathematics to follow this book. Still, you should be familiar with the basics of APIs, SDKs, and Python or one of the other common programming languages. How this book is organized: A road map Generative AI in Action is divided into three main parts, encompassing 13 chapters. Each chapter is crafted to build on previous ones, providing a structured and comprehensive learning experience. The first part, “Foundations of Generative AI,” lays the foundation of generative AI, starting with new use cases and a comprehensive understanding of the basics, including foundational models. It delves into the architecture of LLMs, demonstrating their application across various modalities such as text, images, code, and chat. This section also includes examples to help readers grasp these new AI technologies effectively: Chapter 1 introduces the basics of generative AI, differentiating it from traditional AI and showcasing its potential through various real-world applications. Chapter 2 delves into the architecture and functionality of LLMs, exploring their capabilities and limitations. Chapter 3 covers practical steps to generate text using APIs, including hands-on examples. Chapter 4 shows you how generative AI can create images from text descriptions and understand the underlying models, such as DALL-E. Chapter 5 explores other generative AI applications, such as generating music, code, and 3D models. The book’s second part, “Next steps with generative AI,” focuses on advanced topics crucial for anyone wanting to deploy a GenAI-powered application. This part addresses new architecture patterns and constructs such as prompt engineering, data ABOUT THIS BOOKxx integration, fine-tuning, and model adaptation. It also explores the components of the new GenAI application stack: Chapter 6 is a detailed guide to crafting effective prompts to achieve desired outputs from generative AI models. Chapter 7 explains how to enhance generative AI models by incorporating external data sources. Chapter 8 teaches you how to integrate conversational AI with your enterprise data for more interactive applications. Chapter 9 teaches you techniques for customizing generative AI models to better suit specific use cases. The book’s final section, “Deployment and ethical considerations,” covers best practices for production deployment, scaling strategies, evaluation and benchmarking techniques, and responsible and ethical AI guidelines. These advanced topics are essential for organizations preparing to deploy and utilize generative AI in production at scale: Chapter 10 will help you understand the architectural considerations for developing and deploying generative AI applications. Chapter 11 offers strategies for scaling generative AI models in a production environment. Chapter 12 teaches you how to evaluate and benchmark generative AI models to ensure they meet performance standards. Chapter 13 is a comprehensive guide on the ethical considerations, governance, and safety measures necessary for responsible AI deployment. The book is designed to be read sequentially from cover to cover, as each chapter builds on the concepts introduced in the previous chapters. However, readers already familiar with the basics may focus on specific chapters that address their particular interests or needs. Code samples are included throughout the book to reinforce learning and provide hands-on experience. Running these samples is highly recommended; the code can be found in the book’s GitHub repository. This approach ensures that readers understand the theoretical aspects of generative AI and gain practical skills to implement these technologies effectively. This book focuses on Azure OpenAI and OpenAI, the leading LLM platforms, due to their stability and enterprise readiness. It aims to educate readers on generative AI applications in business, with principles applicable across various LLMs. While it includes diverse LLM examples and open source models, the emphasis is on the Microsoft stack, mainly because it is widely used in the industry and also accessible to the author. About the code This book provides source code for various chapters to enhance the hands-on learning experience. The code is designed to help you practice and apply the concepts discussed in the book. You can download the source code for the relevant chapters of the book. ABOUT THIS BOOK xxi Many examples of source code are contained both in numbered listings and in line with normal text. In both cases, source code is formatted in a fixed-width font like this to separate it from ordinary text. Sometimes, code is also in bold to highlight code that has changed from previous steps in the chapter, such as when a new feature adds to an existing line of code. In many cases, the original source code has been reformatted; we’ve added line breaks and reworked indentation to accommodate the available page space in the book. In rare cases, even this was not enough, and listings include line-continuation markers (➥). Additionally, comments in the source code have often been removed from the listings when the code is described in the text. Code annotations accompany many of the listings, highlighting important concepts. You can get executable snippets of code from the liveBook (online) version of this book at https://livebook.manning.com/book/generative-ai-in-action. The complete code for the examples in the book is available for download from the Manning website at www.manning.com/books/generative-ai-in-action, and from GitHub at https:// github.com/bahree/GenAIBook. You will need the following software and versions to run the provided code: IDE—Visual Studio Code (or similar). Python—Version 3.7.1 or later; we use version 3.11.3 for the book. Package manager—Although technically a package manager is not needed, it would make things much easier to maintain. We use conda for the book, but you can use any package manager. Git—Given we are using GitHub, you need Git installed locally. Docker—Used for containerized deployments and reproducible environments. In the second part of the book, containers are utilized for more advanced use cases. Various SDKs—Used for text and image generation examples, including Azure OpenAI, OpenAI, Gemini, etc. Various other packages—Used for working through different aspects of the chapters. I edited most of the book’s code for clarity and brevity. For example, I left out some things that are not very useful in a printed book, such as exception handling, boilerplate functions, and so forth. The GitHub repository has all these, and the code there is tested and runnable. These tools and libraries are essential for running the examples and exercises provided in the book. Ensure you have the correct versions installed to avoid compatibility issues. Detailed instructions for setting up the environment and dependencies are included in the GitHub code repository, which can be found at https://github.com/ bahree/GenAIBook. liveBook discussion forum Purchase of Generative AI in Action includes free access to liveBook, Manning’s online reading platform. Using liveBook’s exclusive discussion features, you can attach ABOUT THIS BOOKxxii comments to the book globally or to specific sections or paragraphs. It’s a snap to make notes for yourself, ask and answer technical questions, and receive help from the author and other users. To access the forum, go to https://livebook.manning .com/book/generative-ai-in-action/discussion. You can also learn more about Manning's forums and the rules of conduct at https://livebook.manning.com/discussion. Manning’s commitment to our readers is to provide a venue where a meaningful dialogue between individual readers and between readers and the author can take place. It is not a commitment to any specific amount of participation on the part of the author, whose contribution to the forum remains voluntary (and unpaid). We suggest you try asking the author some challenging questions lest his interest stray! The forum and the archives of previous discussions will be accessible from the publisher’s website as long as the book is in print. xxiii about the author AMIT BAHREE is a Principal TPM at Microsoft, where he is part of the engineering team building the next generation of AI products and services for millions of customers using the Azure AI platform. He is also responsible for custom engineering across the platform with key customers, solving complex enterprise scenarios using all forms of AI, including generative AI. A simple geek at heart, Amit has nearly 30 years of experience in technology and product development. He has a strong background in applied research, machine learning, AI, and cloud platforms. He is passionate about creating potent and responsible AI products that transform industries and improve lives. Amit resides in the Seattle area with his wife, daughter, and the sweetest dog, who is not spoilt rotten. 6CHAPTER 1 Introduction to generative AI 1.2 What is generative AI? Generative AI is not a new field of AI, but it has gained more popularity and attention lately. It can generate new content in various outputs—from realistic human faces and writing persuasive text to composing music and developing novel drug compounds. This new AI technique is about replicating existing patterns, imagining new ones, crafting new scenarios, and creating new knowledge. As shown in figure 1.2, generative AI is a subsection of AI that is trained on a vast array of data to learn the underlying patterns and distributions. The magic lies in its potential to generate something novel and original, a task previously believed to be the sole domain of human ingenuity. Machine and deep learning provide the fundamental techniques we need to understand before diving into generative AI. They give us the toolkit to navigate the landscape of AI and understand the processes behind data engineering, model training, and inference. As we progress through this book, we will apply these principles but will not get into the details. Multiple books have been dedicated to both topics, and it would be more prudent for the reader to consult those for details. At its simplest, machine learning (ML) is the scientific discipline focusing on how computers can learn from data. Instead of explicitly programming computers to carry out tasks, in ML, we develop algorithms that can learn from and make predictions or decisions based on data. This data-driven decision-making is applicable to numerous real-world scenarios, ranging from spam filtering in emails to recommendation systems on e-commerce platforms. Deep learning (DL), a subset of ML, takes this concept further. It uses artificial neural networks with several layers. These networks attempt to simulate the behavior of the human brain—albeit in a simplified form—to learn from large amounts of data. While a neural network with a single layer can still make approximate predictions, additional hidden layers can help optimize its accuracy. DL drives many AI applications today and helps execute tasks with improved efficiency, speed, and scale. Artificial intelligence Machine learning Deep learning Generative AI Figure 1.2 Generative AI overview 71.2 What is generative AI? An AI model is a sophisticated algorithmic structure trained on extensive datasets to autonomously perform specific tasks such as text generation, translation, and decision-making. These models learn from data patterns to mimic human cognitive abilities, which enables them to understand and generate natural language. Once trained, developers should recognize that these models can process and analyze data independently, using ML and DL techniques. ML models apply mathematical frameworks to data for predictions, while DL models use neural networks for complex tasks involving unstructured data. In essence, an AI model is a self-sufficient tool that can carry out intelligent tasks based on learned data patterns after training, which are crucial for creating smart applications. Generative AI is an evolution of DL. Many incorrectly assume that ChatGPT is generative AI. ChatGPT is a web application that uses generative AI at its simplest level. The rise and popularity of ChatGPT exposed many folks to generative AI, and the power of the other generative models called large language models (LLMs) is, as the name suggests, related to language. OpenAI trained ChatGPT on diverse internet text to produce a human-like conversation. In addition to ChatGPT, table 1.1 outlines some of the key generative AI models used today; these are grouped by generated AI area types: language, image, and code generation. Table 1.1 Popular generative AI models Name Description Area Generative Pre-trained Transformer (GPT) A large language model developed by OpenAI and trained on a massive dataset of text and code can generate text, translate languages, write various kinds of creative content, and answer your questions informatively. GPT4-Omni (more commonly referred to as GPT-4o) is a multimodal model. At the time of writing, it is the latest version and is a significant upgrade from GPT-4, offering speed, cost, and capability improvements. Language/ multimodal Llama 3 Meta recently released the third version of a natural large language model, open-sourced under a special license. The models come in various sizes and have varying capabilities. Language Claude 3 Anthropic has introduced the Claude 3 model family, which includes Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus. These models offer a range of capabilities, with Opus being the most intelligent. It is capable of complex tasks and exhibits nearhuman comprehension and fluency levels. Like OpenAI’s ChatGPT, Claude can generate text, write code, summarize, and reason, among other things, for a given prompt. Language Cohere Command Cohere offers two models (Command R and Command R+) as part of its Command family. While these LLMMS are optimized for various use cases, Cohere’s newest large language model, Command R+, is optimized for conversational interaction and longcontext tasks. It is designed to be highly performant for complex retrieval-augmented generation (RAG) workflows and multistep tool use. Language 8CHAPTER 1 Introduction to generative AI The following list describes a few areas where generative AI is used today. We expect to see even more innovative and creative applications as generative AI technology develops: Images—This technology creates realistic images of people, objects, and scenes that do not exist in the real world. It is used for various purposes, such as creating virtual worlds for gaming and entertainment, generating realistic product images for e-commerce, and training data for other AI models. Videos—Creates videos that do not exist in the real world. This technology is used for various purposes, such as creating special effects for movies and TV shows, generating training data for other AI models, and creating personalized video content for marketing and advertising. Mistral Mistral Large Language Models are advanced AI models designed for text generation and other language tasks. They have models in different sizes from a collection of open source models (Mistral7B, 8x7B, and 8x22B) and optimized commercial models (Mistral Small, Medium, and Large), each tailored for different reasoning complexities and workloads. Language Gemini Gemini is Google’s new multimodal model that can understand text, images, videos, and audio. It will be available in different sizes (Ultra, Pro, and Nano), each with different capabilities. Language/ multimodal DALL-E Visual AI model developed by OpenAI that can create realistic images from text prompts Image Stable Diffusion Open source image generation model that generates images from a prompt as input. It is primarily used to generate detailed images conditioned on text descriptions and can also be applied to other tasks such as inpainting, outpainting, and generating image-toimage translations. Image Midjourney An image generation model using natural language prompts from a startup called Midjourney, Inc., similar to OpenAI’s DALL-E and Stable Diffusion. Image CodeWhisperer CodeWhisperer is an AWS code-generation model that can generate code in several programming languages, including Python, Java, JavaScript, and TypeScript. Code CodeLlama CodeLlama is a large language model built on Llama 2 and specifically trained on code. It is available in various sizes and supports multiple popular programming languages. Code Codex A large language model is trained specifically on code and used to help with code generation. It supports over a dozen programming languages, including some of the more commonly used, such as C#, Java, Python, JavaScript, SQL, Go, PHP, and Shell, among others. Code Table 1.1 Popular generative AI models (continued) Name Description Area 91.3 What can we generate? Text (language)—This technology creates realistic text, such as news articles, blog posts, and creative writing. It is used for various purposes, such as generating content for websites and social media, creating personalized marketing materials, and creating synthetic data. Text (code)—Generative AI models augment and assist developers when they write code. GitHub’s research found that developers who use its Copilot feature feel 88% more productive and are 96% faster on repetitive tasks. Music—Generative AI models are being used to create original and creative new music. This technology serves various purposes, such as creating music for movies and TV shows, generating personalized playlists, and creating training data for other AI models. We’ll dive into the specifics of how generative AI works in the next chapter, but for now, let’s discuss what can be generated using this technology and how it can help your enterprise. 1.3 What can we generate? When it comes to generating things using generative AI, the sky is the limit. As discussed earlier, we can generate text, images, music, code, voice, and even designs. Before we look at some examples of things that can be generated, it is worth noting that generative AI does not understand the content as humans do. It uses patterns in the data (part of its training set) to generate new, similar data—the quality and relevance of the generated content are directly correlated to the quality and relevance of the training data. 1.3.1 Entities extraction We can use generative AI, specifically a large language model (LLM), to extract entities from text. Entities are pieces of information that are of interest to us. In the past, we would need to use a named entity recognition (NER) model for entity extraction; furthermore, that model would need to have seen the data and be trained as part of its dataset. With LLM models, we can do this without any training, and they are more accurate. While traditional NER methods are effective, they often require manual effort and domain-specific customization. LLMs have significantly reduced this burden, offering a more efficient and often more accurate approach to NER across various domains. A key reason is the Transformer architecture, which we will cover in the next few chapters. This is a great example of traditional AI being more rigid and less flexible than generative AI. Here, we will use OpenAI’s GPT-4 model to extract the first name, company name, location, email, and phone number from the text: Extract the name, company, email, and phone number from the text below: Hello. My name is Amit Bahree. I’m calling from Acme Insurance, Seattle, WA. My colleague mentioned that you are interested in learning about our comprehensive 10 CHAPTER 1 Introduction to generative AI benefits policy. Could you give me a call back at (555) 111-2222 when you get a chance so we can go over the benefits? I can be reached Monday to Friday during normal business hours. If you want, you can also email me at aweomsein- [email protected]m. Thanks, Amit. We can see the entities extracted by the model in the output in figure 1.3. Figure 1.3 Entity extraction 1.3.2 Generating text Generating text in many ways is the bread and butter of generative AI. We can generate different things, such as poems, jokes, and personalized content. In this example, we will use GPT-4 to generate a book launch email for a new book on generative AI. We give the model a prompt with our requirements and some context from which to build. We set up the scene, outline the requirements, and add additional context, such as the text's tone, which will generate the email for us. Given the broad nature of the underlying model (called foundational models, covered in the next chapter), it is also good at understanding grammar and spelling mistakes. For example, we misspelled “pracrtices” and “architectes,” and it understood both: Write a short book launch email for the new generative AI book available at all bookstores. The book covers introduction to generative AI, the different types and use cases, and best pracrtices for enterprises, including how they can incorporate this new AI technology in their business processes. The book is published by Manning and the author is Amit Bahree. The name of the book is Generative AI in Action: Real-World 111.3 What can we generate? Applications and Case Studies of AI for Enterprises. The target audience is developers and enterprise architectes, and the tone is friendly and exciting. 1What should be the subject line of the email? 2What should be the body of the email? We get the following answer, which is rather good: 1Subject Line: “Explore the Power of Generative AI with our New Book: ‘Generative AI in Action’ by Amit Bahree!” 2Email Body: Dear [Recipient's Name], I am thrilled to announce the release of a remarkable new book, “Generative AI in Action: Real-World Applications and Case Studies of AI for Enterprises”, authored by Amit Bahree and published by Manning. Our dynamic business environment is calling for cutting-edge solutions, and Generative AI is just that. It is one of the most exciting advancements in the tech industry and can be a game-changer for your business. Whether you're a developer or an enterprise architect, this book is your ticket to harnessing the power of Generative AI. “Generative AI in Action” unravels the complexities of Generative AI. It provides an accessible introduction and dives deep into its different types, use cases, and how it can revolutionize business processes. Furthermore, it offers the best practices for enterprises looking to incorporate this state-of-the-art technology into their operations. Here's what you can expect from this enlightening read: – A comprehensive introduction to Generative AI. – Exploration of various types of Generative AI and their real-world applications. – Insightful case studies demonstrating how businesses leverage Generative AI. – Valuable guidance on the best practices for implementing Generative AI in your enterprise. We’ve carefully curated this book to make Generative AI accessible to anyone eager to understand and use it. So whether you're new to AI or an experienced professional seeking to stay at the forefront of technological advancements, “Generative AI in Action” is the resource you've been waiting for. We're excited to share that “Generative AI in Action” is now available at all leading bookstores. Grab your copy today and start your journey towards transforming your business with Generative AI. Thank you for your continued support and happy reading! Best Regards, [Your Name] [Your Position] [Your Contact Information] 12 CHAPTER 1 Introduction to generative AI 1.3.3 Generating images We are not restricted to only text—we can generate images using a prompt. To be clear, the model does not find a similar image that satisfies these attributes; rather, it generates something new that does not exist. It does not matter if we ask for something whimsical that will not exist in the real world, such as a dog writing a book with the Taj Mahal in the background. The AI can deconstruct and comprehend each aspect and then use that to create something new, similar to how an artist would. In figure 1.4, we use OpenAI’s DALL-E 3 model to generate the images from the prompt. Generate an image of a dog wearing glasses sitting at a table and authoring a book on AI using a computer. Make it a positive image with the background of the Taj Mahal in the window in the distance at the golden hour. Figure 1.4 Image generation using DALL-E 3 1.3.4 Generating code When thinking about generating code, it is helpful to think of AI not as being able to create fully functioning applications but rather as being able to create some functions and routines. A lot of code is about scaffolding of different runtimes and frameworks and less about the exact business logic. In many of these scenarios, code generation can help improve the developer’s productivity. In the following example, we use 131.3 What can we generate? GPT-3.5 to generate code for a classic “Hello, World!” function. We can give it a prompt such as the following, and it will generate the code for us. Write a hello world equivalent in Python using OpenAI API’s for a developer who is new to using OpenAI and translate the output into French. You get an answer like listing 1.1, including the steps required to start, which is impressive. Of course, this is just an illustrative example to show the model’s power— understanding the context and rules of the request, including the programming language, the software development kit (SDK), packages to use, and, finally, generating code. This code does not follow established best practices (e.g., one should not have their API key in the code). import os from openai import OpenAI gpt_model = "gpt-3.5-turbo" # Replace with your actual OpenAI API key client = OpenAI(api_key='your-api-key') # Generate English text response_english = client.chat.completions.create( model="gpt-3.5-turbo", messages=[ { "role": "user", "content": "Hello, World!" } ], max_tokens=50 ) english_text = response_english.choices[0].message.content.strip() print(english_text) # Translate English text to French response_french = client.chat.completions.create( model="gpt-3.5-turbo", messages=[ { "role": "user", "content": "Translate the following English ➥text to French: " + english_text } ], max_tokens=100 ) Listing 1.1 “Hello, World!” calling OpenAI service 14 CHAPTER 1 Introduction to generative AI # This prints the translation to French print(response_french.choices[0].message.content.strip()) Run the script: python helloworld.py The output is shown in figure 1.5. Note that given the non-deterministic nature of AI, we get a slightly different response each time we run this. In chapter 3, we will see how we can control some of this using different options and nudge the model toward being more deterministic. Some of these options are how large language models (LLMs), a type of generative AI model, work. We will explore LLMs in the next chapter. Figure 1.5 Generated code While we outline code generation, LLMs can be helpful in other aspects related to code, such as explainability, optimization, and readability. We will get into more details about the code later in chapter 5. 1.3.5 Ability to solve logic problems Generative AI's ability to reason is impressive. Due to the emergent behavior of these AI techniques, if we can express the problem domain via a prompt, we can get an answer and a step-by-step thought process of the approach. This resembles how humans think, break down, and approach a problem. For example, we can ask the model to solve a simple math equation, explain the steps, and give us the answer. Solve this problem and think it through step by step. If x = 80 and y = 60 What is x when y = 30? 151.3 What can we generate? Here is the response showing the model working through its steps and the exact thought process it underwent. Achieving something like this that can generalize across a wider domain space with traditional AI would be difficult, if not impossible: Given that x = 80 when y = 60, we can establish a proportion to solve for x when y = 30. In this case, we assume a linear relationship between x and y, that is, x is directly proportional to y. So, when y is halved, x will also be halved. Using the given values: x1/y1 = x2/y2 Where: x1 = 80 y1 = 60 y2 = 30 and we're trying to find x2. We can set up the equation: 80/60 = x2/30 Solving for x2, we get: x2 = (80/60) * 30 This simplifies to: x2 = (4/3) * 30 x2 = 40 So, if y = 30, then x = 40. Traditional AI cannot solve this problem because it has a narrow understanding of the world. It cannot understand the question’s intent to outline a step-by-step process because traditional AI is not generative. It can only follow instructions that are explicitly programmed into it. Generative AI, however, can comprehend the intent of the question because it has a broader understanding of the world and can generate step-by-step processes. Another aspect that allows this is the emergent behavior of generative AI models. This behavior is the ability to outline a step-by-step process. It is not present in any of the individual components of the model but emerges from the interaction of the components. The next chapter will cover emergent behavior in more detail when introducing large language models. 1.3.6 Generating music Similar to how we can use prompts and generate images, we can do the same with music. Music generation is still new compared to text, but there are rapid 22 CHAPTER 1 Introduction to generative AI model such as GPT-4, a big language model, does not make any difference by itself. These advanced AI systems must be implemented and connected to the enterprise’s business lines and processes like any other external software. We will see examples of how to implement this in subsequent chapters. At a high level, there should be few changes from an overall approach; enterprises are still advised to take a thoughtful and strategic approach when incorporating generative AI. The following are a few key considerations—these span various dimensions that most enterprises need to consider, from strategic to business to technical: Crawl, walk, and run. Start small, and do not rush in to do too much too soon. Start with a small pilot project to evaluate, learn, and adapt. This is a complex technology, and it takes time to develop and deploy effective generative AI applications. Do not expect to see results overnight. Define clear objectives and the right use cases. It is important for enterprises to carefully evaluate potential use cases and select those that are most likely to deliver value. The selected use case will guide the choice of AI models, data preparation, and resource allocations. Some generative AI applications are more mature and have a proven record of success, while others are still in their early days. Establish governance policies. Generative AI can generate data, some of which may be sensitive or harmful. Enterprises must establish governance policies to ensure this data is used responsibly and securely. These policies should address problems such as data ownership, privacy, and security. Establish responsible AI and ethical governance. Considering the ethical implications of using generative AI is important. Establish a separate responsible AI and ethical set of policies that reflect the company’s values and that are important to managing its reputation and brand. This includes concerns around bias in AI outputs, the potential misuse of generated content, hallucinations and incorrect details in generated content, and the implications of automating tasks that humans previously performed. A robust AI governance and ethics framework can help manage these risks. Experiment and iterate. Unlike computer science, AI, particularly generative AI, is nondeterministic, and depending on the model parameters and settings, the output can be quite different. As with any AI application, it is essential to take an iterative approach when implementing generative AI. Start with smaller projects, learn from the outcomes, and gradually scale up. This approach helps to manage risk and gain practical experience. Design for failure. Most generative AI models today are commercially available as cloud APIs. As such, they are complex and have a considerable latency compared to more traditional APIs. Enterprises should adhere to cloud best practices and design for failure. They should also factor in best practices of retry mechanics, including exponential backoff policies, caching, security, etc. Expand existing architecture. These new generative AI endpoints are just additional pieces of the overall system. As such, most organizations will want to keep their 231.8 Architecture considerations existing architecture guidance and practices and expand their existing architecture and best practices, rather than starting from scratch. New constructs, such as context windows, tokens, and embeddings, need to be incorporated. Bring your data. One of the main differentiators enterprises have is their proprietary data and associated prompts; therefore, determining how one can utilize their proprietary internal data when using GenAI-powered applications is crucial. This needs to be anchored in the use cases at hand, and if not managed properly, it can get complex quickly, which will be covered in later chapters when we talk about RAG. Manage cost. Generative AI is complex and much more expensive. The cost is typically measured differently (such as in tokens) and not in API calls. Much of this is new and different for enterprises, and the costs can easily get out of hand. Complement traditional AI. In most cases, generative AI would help assist existing investment in traditional AI that enterprises already have. Both sets of technologies are not mutually exclusive but rather support each other. Open-source versus commercial models. Some models are commercially available, such as Azure OpenAI’s GPT models, and some are open source, such as Stable Diffusion. Depending on the use case, it is important to validate which models to use, what the licensing allows, and what legal and regulatory aspects are already covered. 1.8 Architecture considerations Suppose you are an enterprise developer who is seeing all the news on generative AI and the various product announcements from major technology companies. In that case, you might think that for AI, everything has changed. Still, in reality, nothing has changed. From an enterprise perspective, there are new aspects of generative AI that one needs to consider—most, if not all, of these would be things to add to existing architecture best practices and guidance, rather than throwing out anything. We will cover the details later in the book, but new architectural patterns must be accounted for at a high level. We have already touched on many of these, but the key ones are Prompts—We will see how to assess engineering and managing aspects around prompts, including tokens and context windows. Model adaptation—The aim is to make the output better for specific tasks. Integrating generative AI into existing enterprise line-of-business systems—These new AI models alone do not solve a business problem. Design for failure—This aspect is nothing new per se when building missioncritical systems, but many still take shortcuts. Cost and ROI—These generative AI systems are tremendously expensive because the underlying compute is very expensive as well. The costs will come down over time, but they must be consciously planned and designed up front. 24 CHAPTER 1 Introduction to generative AI For example, the cost of GPT-3.5 Turbo from OpenAI came down by 90%, and its quality went up by 90% compared to GPT-3 [5]. Implement policies and approaches for open source (OSS) versus commercial models— Each week, newer models power AI systems and are released. Some are commercial and others are OSS, with different licensing structures. Vendor—There are a few vendors in production that enterprises can use today, but more are coming soon. Today, two of the most mature are OpenAI and Azure OpenAI. The former targets smaller companies and startups, whereas the latter targets enterprises. Google is also releasing its generative AI suite on Google Cloud, and there have been similar announcements from Amazon. In addition, many well-funded startups have announced similar products, such as Anthropic and Mistral. Enterprises need to consider each as a vendor and identify which one they would want to utilize and depend on. 1.9 So your enterprise wants to use generative AI. Now what? Your enterprise has taken a critical step toward using generative AI to drive innovation and efficiency. However, understanding what comes next is crucial to maximizing the benefits and mitigating the risks of this advanced technology. To get started, we will use the example of implementing an Enterprise ChatGPT and outline the steps needed at a high level. Throughout the next few chapters, we will dig into more technical details, including guidance on implementation and best practices. Figure 1.7 shows a high-level overview of what a typical workflow in an enterprise might look like. Figure 1.7 High-level overview of implementing generative AI You should start by setting clear goals for your chatbot. What challenges do you want to address with generative AI? How can it help you the most? This could be anything from creating content for marketing to enhancing customer service with chatbots, Goals • Use cases • Success criteria 1 Resources • People • Software • Hardware 2 Data • Cleaning • Ingestion • Indexing 3 Integrate generative AI • Line of business app • Prompt engineering • Safe and responsible AI 4 Deploy • Test MVP • Deploy to production • Monitor 5 25Summary forecasting for business plans, or even innovating new products or services. In our example, we are building an Enterprise ChatGPT, such as OpenAI’s ChatGPT, but one that is deployed and runs in an enterprise environment, using internal and proprietary data, and only authorized users can access it. Next, we need to ensure that we have the necessary resources available, that is, people with the right competencies, a suitable hardware and software framework, defining indicators of success, and the appropriate governance and ethics principles in place. Then, consider the data. In our example, the enterprise chatbot would need access to relevant, high-quality enterprise data that the user can employ. This data needs to be ingested and indexed to help answer proprietary questions. Before that, the data must be managed properly, ensuring privacy and legal compliance. Remember, the quality of the data fed will influence the output quality. Next, we need to integrate the enterprise chatbot into the line of business applications that address the use case and the problem we are trying to address. As an enterprise, we will also want to address the risks associated with generative AI and implement corporate guidance around safety and responsible AI. Lastly, although we might be ready to deploy in production, implementing generative AI is not a one-time event but a journey. It requires continuous monitoring, testing, and fine-tuning to ensure it works optimally and responsibly. It’s a good idea to start with smaller, manageable projects and gradually scale up as you gain more confidence and expertise in handling this powerful technology. Adopting generative AI is a significant commitment that could transform your enterprise, but it requires careful planning, appropriate resources, ongoing monitoring, and an unwavering focus on ethical considerations. With these in place, your enterprise can reap the numerous benefits of generative AI. Summary Generative AI can be used for multiple use cases, such as entity extraction; generating specific and personalized text, images, code, and music; interpreting text; and solving logical problems. Generative AI use cases can be horizontal across most industries (such as customer services and personalized marketing) or industry specific (such as fraud detection in finance or personalized treatment plans in healthcare). Traditional AI predominantly operates in predefined narrow lanes and can act only in those dimensions, unlike generative AI, which is broader and allows for more flexibility. This chapter outlined an approach and architecture considerations for enterprises to use when adopting and implementing generative AI. 26 Introduction to large language models Large language models (LLMs) are generative AI models that can understand and generate human-like text based on a given input. LLMs are the foundation of many natural language processing (NLP) tasks, such as search, speech-to-text, sentiment analysis, text summarization, and more. In addition, they are general-purpose language models that are pretrained and can be fine-tuned for specific tasks and purposes. This chapter covers An overview of LLMs Key use cases powered by LLMs Foundational models and their effect on AI development New architecture concepts for LLMs, such as prompts, prompt engineering, embeddings, tokens, model parameters, context window, and emergent behavior An overview of small language models Comparison of open source and commercial LLMs 272.1 Overview of foundational models This chapter explores the fascinating world of LLMs and their transformative effect on artificial intelligence (AI). As a significant advancement in AI, LLMs have demonstrated remarkable capabilities in understanding and generating human-like text, thus enabling numerous applications across various industries. Here, we dive into the critical use cases of LLMs, the different types of LLMs, and the concept of foundational models that has revolutionized AI development. The chapter discusses essential LLM concepts, such as prompts, prompt engineering, embeddings, tokens, model parameters, context windows, transformer architecture, and emergent behavior. Finally, we compare open source and commercial LLMs, highlighting their advantages and disadvantages. By the end of this chapter, you will have a comprehensive understanding of LLMs and their implications for AI applications and research. LLMs are built on foundational models; therefore, we will start by outlining what these models are before discussing LLMs in more depth. 2.1 Overview of foundational models Introduced by Stanford researchers in 2021, foundational models have substantially transformed the construction of AI systems. They diverge from task-specific models, shifting to broader, more adaptable models trained on large data volumes. These models can excel in diverse natural language tasks, such as machine translation and question answering, as they learn general language representations from extensive text and code datasets. These representations can then be used to perform various tasks, even tasks they were not explicitly trained on, as shown in figure 2.1. In more technical terms, foundational models utilize established machine learning techniques such as self-supervised learning and transfer learning, enabling them to apply acquired knowledge across various tasks. Developed by means of deep learning, these models employ multilayered artificial neural networks to comprehend complex data patterns; hence, their proficiency with unstructured data such as images, audio, and text. This also extends to 3D signals—data representing 3D attributes that capture spatial dimensions and depth, such as 3D point clouds from LiDAR sensors, 3D medical imaging such as CT scans, or 3D models used in computer graphics and simulations. These can be utilized to make predictions based on 3D data for tasks such as object recognition, scene understanding, and navigation in robotics and autonomous vehicles. NOTE Transfer learning is a machine learning technique in which a model developed for one task is reused as a starting point for a similar task. Instead of starting from scratch, we use the knowledge from the previous task to perform better on the new one. It’s like using knowledge from a previous job to excel at a new but related job. Generative AI and foundational models are closely interlinked. As outlined, foundational models, trained on massive datasets, can be adapted to perform various tasks; this property makes them particularly suitable for generative AI and allows for creating 28 CHAPTER 2 Introduction to large language models new content. The broad knowledge base of these models allows for effective transfer learning, which can be used to generate new, contextually appropriate content across diverse domains. They represent a unified approach, where a single model can generate various outputs, offering state-of-the-art performance owing to their extensive training. Without foundational models as the backbone, there would be no generative AI models. Figure 2.1 Foundational model overview Here are some examples of the common foundation models: GPT (Generative Pre-trained Transformer) Family is an NLP family of models developed by OpenAI. It is a large language model trained on a massive dataset of text and code, which makes it capable of generating text, translating languages, writing creative content, and answering your questions informatively. GPT-4, the latest version at the time of this writing, is also a multimodal model—it can manage both language and images. Codex is a large language model trained specifically on code that is used to help with code generation. It supports over a dozen programming languages, Foundational model Transformer model Text Images Speech Structured data 3D signals Q&A Sentiment analysis Information extraction Image captioning Object recognition Instruction follow Code generation Code understanding Tasks AdaptationTraining Data 292.2 Overview of LLMs including some of the more commonly used, such as C#, Java, Python, JavaScript, SQL, Go, PHP, and Shell, among others. Claude is an LLM built by a startup called Anthropic. Like OpenAI’s ChatGPT, it predicts the next token in a sequence when given a certain prompt and can generate text, write code, summarize, and reason. BERT (Bidirectional Encoder Representations from Transformers) is an NLP model developed by Google. It is a bidirectional model, meaning it can process text in both directions, from left to right and right to left. This feature makes it better at understanding the context of words and phrases. PaLM (Pathway Language Model) and its successor PaLM2 are large multimodal language models developed by Google. The multimodal model can process text, code, and images simultaneously, making it capable of performing a wider range of tasks across those modalities compared to traditional language models operating only in one modality. Gemini is Google’s latest AI model, capable of understanding text, images, videos, and audio. It’s a multimodal model described as being able to complete complex tasks in math, physics, and other areas, as well as understanding and generating high-quality code in various programming languages. Gemini was built from the ground up to be multimodal, meaning it can generalize and seamlessly understand, operate across, and combine different types of information. It’s also the new umbrella name for all of Google’s AI tools, replacing Google Bard and Duet AI, and is considered a successor to the PaLM model. Once a foundational model is trained, it can be adapted to a wide range of downstream tasks by fine-tuning its parameters. Fine-tuning involves adjusting the model’s parameters to optimize the model for a specific task. It can be done using a small amount of labeled data. By fine-tuning these models for specific tasks or domains, we use their general understanding of language and supplement it with task-specific knowledge. The benefits of this approach include time and resource efficiency, coupled with remarkable versatility. We can also adapt a model via Prompt engineering, which we’ll discuss later in this chapter. Now that we know more about foundational models, let’s explore LLMs. 2.2 Overview of LLMs LLMs represent a significant advancement in AI. They are trained on a vast amount of text data, such as books, articles, and websites, to learn patterns in human language. They are also hard to develop and maintain, as they require lots of data, computing, and engineering resources. OpenAI’s ChatGPT is an example of an LLM—it generates human-like text by predicting the probability of a word considering the words already used in the text. The model learns to generate coherent and contextually relevant sentences by adjusting its internal parameters to minimize the difference between its predictions 30 CHAPTER 2 Introduction to large language models and the actual outcomes in the training data. When generating text, the model chooses the word with the highest probability as its subsequent output and then repeats the process for the next word. LLMs are foundational models adapted for natural language processing and language generation tasks. These LLMs are general-purpose and can handle tasks without task-specific training data. As briefly described in the previous chapter, given the right prompt, they can answer questions, write essays, summarize texts, translate languages, and even generate code. LLMs can be applied to many applications across different industries, as outlined in chapter 1—from summarization to classification, Q&A chatbots, content generation, data analysis, entity extraction, and more. Before we get into more details of LLMs, let us look at the Transformer architecture, which makes these foundational models possible. 2.3 Transformer architecture Transformers are the bedrock of foundational models and are responsible for their remarkable language understanding capabilities. The Transformer model was first introduced in the paper “Attention Is All You Need” by Vaswani et al. in 2017 [1]. Since then, Transformer-based models have become state-of-the-art for many tasks. GPT and BERT are examples of Transformer-based models, and the “T” in GPT stands for Transformers. At their core, Transformers use a mechanism known as attention (specifically selfattention), which allows the model to consider the entire context of a sentence, considering all words simultaneously rather than processing the sentence word by word. This approach is more efficient and can improve the results of many NLP tasks. The strength of this approach is that it captures dependencies regardless of their position in the text, which is an essential factor in language understanding. This is key for tasks such as machine translation and text summarization, where the meaning of a sentence can depend on terms that are several words apart. Transformers can parallelize their computations, which makes them much faster to train than other types of neural networks. This mechanism enables the model to pay attention to the most relevant parts of the task input. In the context of generative AI, a transformer model would take an input (such as a prompt) and generate an output (such as the next word or the completion of the sentence) by weighing the importance of each part of the input in generating the output. For example, in the sentence “The cat sat on the...,” a Transformer model would likely give much weight to the word “cat” when determining that the likely next word might be “mat.” These models exhibit generative properties by predicting the next item in a sequence—the next word in a sentence or the next note in a melody. We explore this more in the next chapter. Transformer models are usually very large, requiring significant computational resources to train and use. Using a car analogy, think of Transformer models as 312.5 Types of LLMs supercharged engines that need much power to run but do amazing things. Think of them as the next step after models such as ResNET 50, which is used for recognizing images. While ResNET 50 is like a car with 50 gears, OpenAI’s GPT-3 is like a megatruck with 96 gears and extra features. Because of their advanced capabilities, these models are a top pick for creating intelligent AI outputs. LLMs use transformers, which are composed of an encoder and a decoder. The encoder processes the input text (i.e., the prompt) and generates a sequence of hidden states that represent the meaning of the input text. The decoder uses these hidden states to generate the output text. These encoders and decoders form one layer, similar to a mini-brain. Multiple layers can be stacked one upon another. As outlined earlier, GPT3 is a decoder-only model with 96 layers. 2.4 Training cutoff In the context of foundational models, the training cutoff refers to the point at which the model’s training ends, that is, the time until the data used to train the model was collected. In the case of AI models developed by OpenAI, such as GPT-3 or GPT-4, the training cutoff is when the model was last trained on new data. This cutoff is important because after this point, the model is not aware of any events, advancements, new concepts, or changes in language usage. For example, the training data cutoff for the GPT-3.5 Turbo was in September 2021, GPT-4 Turbo in April 2023, and GPT-4o in October 2023, meaning the model does not know about real-world events or advancements in various fields beyond that point. The key point is that while these models can generate text based on the data they were trained on, they do not learn or update their knowledge after the training cutoff. They cannot access or retrieve real-time information from the internet or any external database. Their responses are generated purely based on patterns they have learned during their training period. NOTE The recent announcement that the premium versions of ChatGPT will have access to the internet via the Bing plugin doesn’t mean that the model has more up-to-date information. This uses a pattern called RAG (retrievalaugmented generation), which will be covered later in chapter 7. 2.5 Types of LLMs As shown in table 2.1, there are three categories of LLMs. When we talk about LLMs, having the context is crucial, and it might not be evident in some cases. This is of great importance, as the paths we can go down when using the models aren’t interchangeable, and picking the right type depends on the use case one tries to solve. Furthermore, there is also a dependency on how effectively one can adapt the models to specific use cases. 38 CHAPTER 2 Introduction to large language models 2.8 Key concepts of LLMs This section describes the architecture of a typical LLM implementation. Figure 2.3 shows the abstract structure of a common LLM implementation at a high level; it follows this process whenever we use an LLM such as OpenAI’s GPT. Figure 2.3 Conceptual architecture of an LLM The model starts with the input text—the prompt. It is first converted into a sequence of tokens using tokenization. Each token is then converted into a numerical vector via a process called embedding, which acts as the encoder input. (continued) Azure OpenAI offers additional enterprise-ready features, such as data privacy, customer-managed keys, encryption at rest, private networking, regional availability, and responsible AI content filtering. These features can be important for businesses that need to comply with specific security or regulatory requirements. Consequently, the APIs between the two are similar but not the same. However, the underlying models are the same, and Azure OpenAI has a deployment that incorporates these additional features that most enterprises require. Input text (prompt) Token Embedding Encoder Decoder ………………….. ………………….. ………………….. Generated text (completion) Numerical representation Needed for scenarios such as “Bring your own data,” search, etc. LLM ………………….. ………………….. ………………….. ………………….. ………………….. …………………. 392.8 Key concepts of LLMs The encoder processes the input sequence and generates a sequence of hidden states. These hidden states are then fed into the decoder with a start token. The decoder generates the output sequence one token at a time by predicting the next token based on the previous tokens and hidden states. Once all the layers have processed the information, the model predicts the next token in the learned sequence. This outcome is converted back to the text, and we see the response. This process runs in an iterative loop and occurs for each new token generated, thus creating a coherent text output. The final text that the model generates is an emergent property of this layered, iterative process. The final output sequence is also called a completion. Let’s examine each of these LLM aspects in more detail. 2.8.1 Prompts A prompt is how we “talk” to these models. A prompt is just text describing the task we are trying to achieve using natural language. The output of these models is also text. The ability to express our intention in this manner (natural language) instead of conforming to the input restrictions of a machine makes prompts powerful. Crafting or designing the text in the prompt is akin to programming the model and creating a new paradigm called prompt engineering, which will be covered later in the book. Prompts allow us to capture a broader set of intents and emotions using the same paradigm. In the following two examples (generated by AI), a prompt asks for a Haiku outlining how great LLMs are: one haiku is happy, and the other is sad. Getting this level of emotion and expressiveness on the same task was impossible. Write a happy haiku on how great LLMs are. AI-generated response LLMs are great They can chat and generate Creative and smart Write a sad haiku on how great LLMs are. AI-generated response LLMs are great But they cannot feel or love Only simulate NOTE LLMs do not understand text like humans do. They are statistical models that learn patterns in data and use those patterns to generate output. Despite their impressive output, these models don’t comprehend the content generated in the same sense as we do—including beliefs, desires, consciousness, and right and wrong. They are just extremely good at predicting the next word in a sequence of text based on patterns they’ve seen millions of times. 40 CHAPTER 2 Introduction to large language models 2.8.2 Tokens Tokens are the basic units of text that an LLM uses to process both the request and the response, that is, to understand and generate text. Tokenization is the process of converting text into a sequence of smaller units called tokens. When using LLMs, we use tokens to converse with these models, which is one of the most fundamental elements of understanding LLMs. Tokens are the new currency when incorporating LLMs into your application or solutions. They directly correlate with the cost of running models, both in terms of money and of the experience with latency and throughput. The more tokens, the more processing the model must do. This means more computational resources are required for the model, which means lower performance and higher latency. LLMs convert the text into tokens before processing. Depending on the tokenization algorithm, they can be individual characters, words, sub-words, or even larger linguistic units. A rough rule of thumb is that one token is approximately four characters or 0.75 words for English text. For most LLMs today, the token size that they support includes both the input prompt and the response. Let’s illustrate this through an example. Figure 2.4 shows how the sentence “I have a white dog named Champ” gets tokenized (using OpenAI’s tokenizer in this case). Each block represents a different token. In this example, we use eight tokens. Figure 2.4 Tokenizer example LLMs generate text by predicting the next word or symbol (token) most likely to follow a given sequence of words or symbols (tokens) they use as input, that is, the prompt. We show a visual representation of this in figure 2.5, where the list of tokens on the right shows the highest probability of tokens following the prompt “The dog sat on.” We can influence some of this probability of tokens using a few parameters we will see later in the chapter. Suppose we have a sequence of tokens with a length of n. Utilizing these n tokens as the context, we generate the subsequent token, n + 1. This newly predicted token is then appended to the original sequence of tokens, thereby expanding the context. Consequently, the expanded context window for generating token n + 2 becomes Tokenization I have a white dog named Champ . 12345678 412.8 Key concepts of LLMs n + (n + 1). This process is repeated in a continuous loop until a predetermined stop condition, such as a specific sequence or a size limit for the tokens, is reached. For example, if we have a sentence, “Hawaiian pizza is my favorite,” the probability distribution of the next word we see is shown in figure 2.6. The most likely word is “type,” finishing the sentence “Hawaiian pizza is my favorite type.” Figure 2.6 Next token probability distribution If you run this example again, you will get a probability different from the one shown here. This is because most AI is nondeterministic, specifically in the case of LLMs. Simultaneously, it might predict one token, and it is probably being looked at across all the possible tokens that the model has learned in the training phase. We also use two examples that outline how one token changes the distribution dramatically (changing one word from “the” to “a”). Figure 2.7 shows that the most LLM Input token vector Vector representation of next output token the mat pad … … … Highest probability Second highest probability Less likely Next word ……… ……… The dog sat on ……… ……… ……… ……… ……… ……… ……… ……… Figure 2.5 LLM—next token predictor Hawaiian pizza is my favorite 42 CHAPTER 2 Introduction to large language models probable next token is “mat” at 41% probability. We also see a list of the other tokens and their probabilistic distributions. Figure 2.7 Example 1 However, changing one token from “the” to “a” dramatically changes the next distribution set, with the mat jumping up 30 points to a probability of nearly 75%, as shown in figure 2.8. Figure 2.8 Example 2 Some settings related to LLMs are important and can change how the model behaves and generates text. These settings are the model configurations and can be changed via an API, GUI, or both. We cover model configurations in more detail later in the chapter. 2.8.3 Counting tokens Many developers will probably be new to tracking tokens when using LLM, especially in an enterprise setting. However, counting tokens is important for several reasons: Memory limitations—LLMs can process a maximum number of tokens in a single pass. This is due to the memory limitations of their architecture, often defined by their context window (another concept we discuss later in this chapter). For example, OpenAI’s latest GPT-4o model has a content window of 128K, and 432.8 Key concepts of LLMs Google’s latest Gemini 1.5 Pro has a context window of 1M tokens. GPT3.5Turbo, another OpenAI model, has two models supporting 8K and 16K token lengths. There is research ongoing to see how to solve this, such as LongNet [6] from Microsoft Research, which shows how to scale to 1B context windows. It is important to point out that this is still an active research area and has not been productized yet. Cost—When thinking about cost, there are two dimensions: the computational costs in terms of latency, memory, and the overall experience, and the actual cost in terms of money. For each call, the computational resources required for processing tokens directly correlate to the tokens’ length. As the token length increases, it requires more processing time, leading to more computational requirements (specifically memory and GPUs) and higher latency. This also means increased costs for using the LLMs. AI quality—The quality of a model’s output depends on the number of tokens it is asked to generate or process. If the text is too short, the model might not have enough context to provide a good answer. Conversely, if the text is too long, the model might lose coherence in its response. We will touch on the notion of good versus poor as part of prompt engineering later in chapter 6. For many enterprises, cost and performance are key factors in deciding whether to use tokens. Generally speaking, smaller models are more cost-effective and efficient than bigger ones. Listing 2.1 shows a simple way to calculate the number of tokens. In this example, we use an open source library called tiktoken, released by OpenAI. This tokenizer library implements a byte-pair encoding (BPE) algorithm. These tokenizers are designed with their respective LLMs, ensuring efficient tokenization and optimal performance during pretraining and fine-tuning processes. If you use one of the OpenAI models, you must use this tokenizer; many other transformer models also use it. If needed, you can install the tiktoken library using pip install tiktoken import tiktoken as tk def count_tokens(string: str, encoding_name: str) -> int: # Get the encoding encoding = tk.get_encoding(encoding_name) # Encode the string encoded_string = encoding.encode(string) # Count the number of tokens num_tokens = len(encoded_string) return num_tokens # Define the input string prompt = “I have a white dog named Champ” Listing 2.1 Counting tokens for GPT The encoding specifies how the text is converted into tokens. 44 CHAPTER 2 Introduction to large language models # Display the number of tokens in the String print(“Number of tokens:” , count_tokens(prompt, “cl100k_base”)) Running this code, as expected, gives us the following output: $ python countingtokens.py Number of tokens: 7 NOTE Byte-pair encoding (BPE) is a compression algorithm widely used in NLP tasks, such as text classification, text generation, and machine translation. One of the BPE advantages is that it is reversible and lossless, so we can get the original text. BPE works on any text that the tokenizer’s training data hasn’t seen, and it compresses the text, resulting in shorter token sequences than the original text. BPE also helps generalize repeating patterns in a language and provides a better understanding of grammar. For example, the gerund -ing form is quite common in English (swimming, running, debugging, etc.). BPE will split it into different tokens, so “swim” and “-ing” in swimming become two tokens and generalize better. If we are not sure of the name of the encoding to use, instead of the function get_ encoding(), we can use the encoding_for_model()function. This takes the name of the model we want to use and utilizes the corresponding encoding, such as encoding = tiktoken.encoding_for_model('gpt-4'). For OpenAI, table 2.3 shows different supported encodings. Listing 2.2 shows how to use different encodings and how to get the original text from the tokens. We should understand this as a basic construct for now, but it is useful for more advanced use cases such as caching and chunking text—aspects that we cover later in the book. import tiktoken as tk def get_tokens(string: str, encoding_name: str) -> str: # Get the encoding encoding = tk.get_encoding(encoding_name) # Encode the string return encoding.encode(string) Table 2.3 OpenAI encodings Encoding OpenAI model cl100k_base gpt-4, gpt-3.5-turbo, gpt-35-turbo, text-embedding-ada-002 p50k_base Codex models, text-davinci-002, text-davinci-003 r50k_base GPT-3 models (davinci, curie, babage, ada) Listing 2.2 Tokens 452.8 Key concepts of LLMs def get_string(tokens: str, encoding_name: str) -> str: # Get the encoding encoding = tk.get_encoding(encoding_name) # Decode the tokens return encoding.decode(tokens) # Define the input string prompt = “I have a white dog named Champ.” # Display the tokens print(“cl100k_base Tokens:” , get_tokens(prompt, “cl100k_base”)) print(“ p50k_base Tokens:” , get_tokens(prompt, “p50k_base”)) print(“ r50k_base Tokens:” , get_tokens(prompt, “r50k_base”)) print(“Original String:” , get_string([40, 617, 264, 4251, 5679, 7086, 56690, 13], “cl100k_base”)) $ python encodings.py cl100k_base Tokens: [40, 617, 264, 4251, 5679, 7086, 56690, 13] p50k_base Tokens: [40, 423, 257, 2330, 3290, 3706, 29260, 13] r50k_base Tokens: [40, 423, 257, 2330, 3290, 3706, 29260, 13] Original String: I have a white dog named Champ. In addition to the tiktoken library we have been using in the examples, there are a few other popular tokenizers. Remember that each tokenizer is designed for the corresponding LLM and cannot be interchanged: WordPiece—Used by the BERT model from Google, it splits text into smaller units based on the most frequent word pieces, allowing for efficient representation of rare or out-of-vocabulary words. SentencePiece—Meta’s RoBERTa model (Robustly Optimized BERT) uses the model. It combines WordPiece and BPE approaches into a single languageagnostic framework, allowing for more flexibility. T5 tokenizer—Based on SentencePiece, it is used by Google’s T5 model (Text-toText Transfer Transformer). XLM tokenizer—This is used in Meta’s XLM (Cross-lingual Language Model) and implements a BPE method with learned embeddings (BPEmb). It is designed to handle multilingual text and support cross-lingual transfer learning. 2.8.4 Embeddings Embeddings are powerful machine-learning tools for large inputs representing words. They capture semantic similarities in a vector space (i.e., a collection of vectors, as shown in figure 2.9), allowing us to determine if two text chunks represent the same meaning. By providing a similarity score, embeddings can help us better understand the relationships between different pieces of text. The idea behind embeddings is that words with similar meanings should have similar vector representations, as measured by their distances. Vectors with smaller distances between them suggest they are highly related, and those with longer distances 46 CHAPTER 2 Introduction to large language models suggest low relatedness. There are a few ways to measure similarities; we will cover these later in chapter 7. These vectors are learned during training and are used to capture the meaning of words or phrases. AI algorithms can easily utilize these vectors of floating-point numbers. Figure 2.9 Embeddings For example, the word “cat” might be represented by a vector as [0.2, 0.3, -0.1], while the word “dog” might be represented as [0.4, 0.1, 0.2]. These vectors can then be used as input to machine learning models for tasks such as text classification, sentiment analysis, and machine translation. Embeddings are learned when the model is trained on a large corpus of text data. The idea is to capture the meaning of words or phrases based on their context in the training data. Depending on the task, there are several algorithms for creating embeddings: Similarity embeddings are good at capturing semantic similarity between two or more pieces of text. Text search embeddings measure whether long documents are relevant to a short query. Code search embeddings are useful for embedding code snippets and natural language search queries. NOTE Embeddings created by one method cannot be understood by another. In other words, if you create an embedding using OpenAI’s API, embeddings of another provider will not understand the vectors created, and vice versa. Listing 2.3 shows how to get an embedding (from OpenAI in this example). We define a function called get_embedding() that takes a string for which we need to create embeddings as a parameter. The function uses OpenAI’s API to generate an embedding for the input text using the text-embedding-ada-002 model. The embedding is returned as a list of floating-point numbers. import os from openai import OpenAI client = OpenAI(api_key=’your-API-key’) Listing 2.3 Getting an embedding in OpenAI Embedding model I have a white dog. Prompt 0.00608 0.01417 ….. 0.02123 Prompt as vector 472.8 Key concepts of LLMs def get_embedding(text): response = client.embeddings.create( model="text-embedding-ada-002", input=text) return response.data[0].embedding embeddings = get_embedding("I have a white dog named Champ.") print("Embedding Length:", len(embeddings)) print("Embedding:", embeddings[:5]) The vector space resulting from the embedding isn’t a one-to-one mapping to the tokens but can be a lot more. The output of the previous examples is shown next. For brevity, we only show the first five items in the list: print("Embedding Length:", len(embeddings)) print("Embedding:", embeddings[:5]) 2.8.5 Model configuration Most LLMs expose some configuration settings to the user, allowing one to tweak how the model operates and its behavior to some extent. While a few parameters would change depending on the model implementation, the three key configurations are temperature, top probability (top_p), and max response. Note that some implementations might have a different name but mean the same thing. The OpenAI implementation of GPT calls the maximum response as max tokens. Let us explore these in a little more detail. MAX RESPONSE The parameter known as max response essentially defines the upper limit for the text length that the model generates. This means that once the model hits this predetermined length, it halts text generation, regardless of whether it is mid-word or midsentence. It’s crucial to grasp this configuration because there is a size limit to the tokens most models can process. Increasing this size corresponds to heightened computational demands, leading to increased latency and cost. TEMPERATURE When generating text, as with any foundational model, inherent randomness yields a different output each time we call the model. Temperature is one of the most important settings for controlling the degree of the model’s randomness. Typically, this is a value from 0 to 1, with 0 representing a more accurate and predictable output. In contrast, setting a 1 makes the output more diverse and random, as shown in figure 2.10. TOP PROBABILITY (TOP_P) The top probability (top_p) parameter (also known as nucleus sampling) is a setting in language model APIs that steers the randomness of the text-generation process. This parameter allows one to fine-tune the balance between creativity and reliability in the text that the model generates. It defines a threshold probability; only words with probabilities above this threshold are considered when the model generates text. When a language model generates text, it predicts the probability of each word being 54 CHAPTER 2 Introduction to large language models // XXX: fails when a and b refer to same memory location # define XOR_SWAP(a,b) do\ {\ a ^= b;\ b ^= a;\ a ^= b;\ } while (0) // walk inwards from both ends of the String, // swapping until we get to the middle while (str < end) { XOR_SWAP(*str, *end); str++; end--; } # undef XOR_SWAP } } The first step would be to ask our LLM to convert this code to Python using a prompt similar to the next one. We are omitting the C function (listing 2.4) again for brevity. The code is part of the prompt. Convert this C function to Python: #include <string.h> // for strlen() // reverse the given null-terminated String in place void inplace_reverse(char * str) { … } Listing 2.5 shows the output from our model with the code converted to Python. It does a good job of understanding what the C function is doing and outlines Python’s different constraints and properties in the response. Note that we did not give it any hints as to what the function is doing. It deciphered the while loop itself, showing an example of emergent behavior. This function will return the reversed string rather than modify the original string because Python strings are immutable. In other words, once a string is created in Python, it cannot be changed. Hence, the function inplace_reverse is a bit of a misnomer in this Python context. You could consider renaming it to something like reverse_string. def inplace_reverse(str): if str: # Convert the String to a list since strings are immutable in Python lst = list(str) Listing 2.5 Converted Python function for string reversal 552.8 Key concepts of LLMs # Initialize two pointers at the start and end of the list start = 0 end = len(lst) - 1 # Reverse the list in place while start < end: # XOR_SWAP equivalent in Python lst[start], lst[end] = lst[end], lst[start] # Move the pointers towards the center of the list start += 1 end -= 1 # Convert the list back into a string return ''.join(lst) return str We can simplify this further by asking the model the following: Can you further simplify and optimize the Python code? Sure! One of the simplest ways to reverse a string in Python uses slicing. Here’s a function that will return the reverse of a string: def reverse_string(s): return s[::-1] The [::-1] is a slice that starts at the end of the string and moves backward. However, if you want to emulate the same logic as the C program, that is, reversing a string in place using two pointers, here is a more Pythonic way of achieving this: def reverse_string(s): lst = list(s) lst.reverse() return ''.join(lst) In this version, we convert the string to a list, reverse the list in place using the built-in reverse function, and then join the list back into a string. The reverse function in Python uses a similar two-pointer approach, but it’s all under the hood, so the code looks much simpler. The point the response calls out on a similar two-pointer approach demonstrates the emergent property’s power. We never explicitly called this out anywhere in the conversation. In this chapter, we looked at the core concepts that make LLMs work and what we need to be mindful of when implementing these new generative AI techniques in the enterprise. We cover new concepts such as prompts, prompt engineering, embeddings, tokens, model parameters, context windows, transformer architecture, and emergent behavior. Now that we have a broad understanding of LLMs and their implications in the next chapter, let us put this into action and work through an API to generate text, both as completions where we provide instructions via a prompt and as 56 CHAPTER 2 Introduction to large language models chat where we can have a dialogue, and in that, generate text, in the context of the conversation. Summary Large language models (LLMs) represent a major advancement in AI. They are trained on vast amounts of text data to learn patterns in human language. LLMs are general-purpose and can handle tasks without task-specific training data, such as answering questions, writing essays, summarizing texts, translating languages, and generating code. Key LLM use cases include summarization, classification, Q&A/chatbots, content generation, data analysis, translation and localization, process automation, research and development, sentiment analysis, and entity extraction. Types of LLMs include base, instruction-based, and fine-tuned LLM. Each has pros and cons and is powered by foundational models. Foundational models are large AI models trained on vast quantities of data at a massive scale, resulting in models that can be adapted to a wide range of downstream tasks. Some key LLM concepts include prompts, prompt engineering, embeddings, tokens, model parameters, context windows, transformer architecture, and emergent behavior. Open source and commercial LLMs have advantages and disadvantages, with commercial models typically offering state-of-the-art performance and open source models providing more flexibility for customization and integration. Small language models (SLMs) are a new emerging trend of lightweight generative AI models that produce text, summarize documents, translate languages, and answer questions. In some cases, they offer capabilities similar to those of larger models. 57 Working through an API: Generating text We have seen that large language models (LLMs) provide a powerful suite of machine learning tools specifically designed to enhance natural language understanding and generation. OpenAI features two notable APIs: the completion and the chat completion APIs. These APIs, unique in their dynamic and effective This chapter covers Generative AI models and their categorization based on specific applications The process of listing available models, understanding their capabilities, and choosing the appropriate ones The completion API and chat completion API offered by OpenAI Advanced options for completion and chat completion APIs that help us steer the model and hence control the generation The importance of managing tokens in a conversation for improved user experience and cost-effectiveness 58 CHAPTER 3 Working through an API: Generating text text-generation capabilities, resemble human output. In addition, they offer developers exclusive opportunities to craft various applications, from chatbots to writing assistants. OpenAI was the first to introduce the pattern of completion and chat completion APIs, which now embody almost all implementations, especially when companies want to build generative-AI-powered tools and products. The completion API by OpenAI is an advanced tool that generates contextually appropriate and coherent text to complete user prompts. Conversely, the chat completion API was designed to emulate an interaction with a machine learning model, preserving the context of a conversation across multiple exchanges, which makes it suitable for interactive applications. Chapter 3 establishes the groundwork for scaling enterprises. These APIs can significantly accelerate the development of intelligent applications, thereby reducing the time to value. We’ll mostly use OpenAI and Azure OpenAI as illustrative examples, often interchangeably. The code models remain consistent, and the APIs are largely similar. Many enterprises may gravitate toward Azure OpenAI because of the control it offers, while others might favor OpenAI. It is important to note that we assume here that an Azure OpenAI instance has already been deployed as part of your Azure subscription, and we will be referencing it in the context of our examples. This chapter outlines the basics of the completion and the chat completion APIs, including how they differ and when to use each. We will see how to implement them in an application and how we can steer the model generation and its randomness. We’ll also see how to manage tokens, which are key operation considerations when deploying to production. These are the fundamental aspects required to build on for a mission-critical application. But first, let’s start by understanding the different model categories and their advantages. 3.1 Model categories Generative AI models can be classified into various categories based on their specific applications, and each category includes different types of models. We start our discussion by understanding the different classifications of models within generative AI. This understanding will help us identify the range of models available and choose the most appropriate one for a given situation. The availability of different types and models may vary, depending on the API in use. For example, Azure OpenAI and OpenAI provide different versions of LLMs. Some versions might be phased out, some could be limited, and others could be exclusive to a certain organization. Different models have unique features and capabilities, directly affecting their cost and computational requirements. Thus, choosing the right model for each use case is critical. In conventional computer science, the idea that bigger is better has often been applied to memory, storage, CPUs, or bandwidth. However, in the case of LLMs, this principle is not always applicable. OpenAI provides a host of models categorized, as shown in table 3.1. Note that these are the same for both OpenAI and Azure OpenAI, as the underlying models are identical. 593.1 Model categories Each model category contains variations that are further distinguished by certain features such as token size. As discussed in the previous chapter, token size determines a model’s context window, which defines the amount of input and output it can process. For instance, the original GPT-3 models had a maximum token size of 2K. GPT-3.5 Turbo, a subset of models within the GPT-3.5 category, has two versions—one with a token size of 4K and another with a token size of 16K. These are double and quadruple the token size of the original GPT-3 models. Table 3.2 outlines the more popular models and their capabilities. Table 3.1 OpenAI model categories Model category Description GPT-4 The newest and most powerful version is a set of multimodal models. GPT-4 is trained on a larger dataset with more parameters, making it even more capable. It can perform tasks that are out of reach for the previous models. There are various models in the GPT-4 family—GPT-4.0, GPT-4 Turbo, and the latest GPT-4o (omni), a multimodal model and the most powerful in the family at the time of publication. GPT-3.5 A set of models that improve on GPT-3 and can understand and generate natural language or code. When unsure, these should be the default models for most enterprises. DALL.E A model that can generate images when given a prompt Whisper A model that is used for speech-to-text, converting audio into text Embeddings A set of models to convert text into its numerical form GPT-3 (Legacy) A set of models that can generate and understand natural language. These were the original set of models that are now considered legacy. In most cases, we would want to start with one of the newer models, 3.5 or 4.0, which derive from GPT-3. Table 3.2 Model descriptions and capabilities Model Capabilities Ada (legacy) Simple classification, parsing, and formatting of text. This model is part of the GPT-3 legacy. Babbage (legacy) Semantic search ranking, medium complex classification. This model is part of the GPT-3 legacy. Curie (legacy) Answering questions, highly complex classification. This model is part of the GPT-3 legacy. Davinci (legacy) Summarization, generating creative content. This model is part of the GPT-3 legacy. Cushman-Codex (legacy) A descendant of the GPT-3 series, trained in natural language and billions of lines of code. It is the most capable in Python and proficient in over a dozen other programming languages. Davinci-Codex A more capable model of Cushman-codex 60 CHAPTER 3 Working through an API: Generating text Note that the mentioned legacy models are still available and work as intended. However, the newer models are better, having more mindshare and longer support. Most should start with GPT-3.5 Turbo as the default model and use GPT-4 on a case-by-case basis. Sometimes, even a smaller, older model, such as the GPT-3 Curie, is good. This provides the right balance between the model’s capability, cost, and overall performance. In the early days of generative AI, all the models were available only to some. These will vary by company, region, and in the case of Azure, your subscription type, among other things. We have to list the models and their capabilities that are available for us to use. However, before listing models, let us see the dependencies required to get things working. 3.1.1 Dependencies In this section, we call out the run time dependencies and configurations needed at a high level. To get things working, we need at least the following items: Development IDE—We use Visual Studio Code for our examples, but you can use anything you are comfortable with. Python—We use v3.11.3 in this book, but you can use any version as long as it is v3.7.1 or later. The installation instructions are available at https:// www.python.org/ if you need to install Python. OpenAI Python libraries—We use Python libraries for most of the code and the demos. The OpenAI Python library can be a simple installation in conda, using conda install -c conda-forge openai. If you are using pip, use pip install --upgrade openai. There are also software development kits (SDKs) for specific languages if you prefer to use those instead of Python packages. Azure Subscription or OpenAI API access—We use OpenAI’s endpoint and the Azure OpenAI (AOAI) endpoint interchangeably; in most cases, either option will work. Given the emphasis on enterprises for this book, we tend to lean toward using the Azure OpenAI service: GPT3.5-Turbo The most capable GPT-3.5 model optimized for chat use cases is 90% cheaper and more effective than GPT-3 Davinci. GPT-4, GPT-4 Turbo More capable than any GPT-3.5 model. It is able to do more complex tasks and is optimized for chat models. GPT-4o The latest GPT-4o model is more capable than the GPT-4 and GPT-4 Turbo, but it is also twice as fast and 50% cheaper. text-embedding-ada-002, text-embedding-ada-003 This new embedding model replaces five separate models for text search, similarity, and code search, outperforming them at most tasks; furthermore, it is 99.8% cheaper. Table 3.2 Model descriptions and capabilities (continued) Model Capabilities 613.1 Model categories – To use the library with Azure endpoints, we need the api_key. – We also need to set the api_type, api_base, and api_version properties. The api_type must be set to azure, the api_base points to the endpoint that we deploy, and the corresponding version of the API is specified via api_version. – Azure OpenAI uses 'engine' as the parameter to specify the model’s name. When deploying the model in your Azure subscription, this name needs to be set to your chosen name. For example, figure 3.1 is a screenshot of the deployments in one subscription. OpenAI, however, uses the parameter model to specify the model’s name. These model names are standard as they release them. You can find more details on Azure OpenAI and OpenAI at https://mng.bz/yoYd and https://platform.openai.com/docs/. NOTE The GitHub code repository accompanying the book (https://bit .ly/GenAIBook) has the details of the code, including dependencies and instructions. Hardcoding the endpoint and key is not an advisable practice. There are multiple methods to accomplish this task, one of which includes using environment variables. We demonstrate this method in the steps that follow. Other alternatives could be fetching them from secret stores or environment files. For the sake of simplicity, we will stick to environment variables in this guide. However, you are encouraged to adhere to your enterprise’s best practices and recommendations. Setting up the environment variables can be achieved through the following commands. For Windows, these are setx AOAI_KEY "your-openai-key" setx AOAI_ENDPOINT "your-openai-endpoint" NOTE You may need to restart your terminal to read the new variables. On Linux/Mac, we have export AOAI_ENDPOINT=your-openai-endpoint export AOAI_KEY=your-openaikey Bash uses echo export AOAI_KEY="YOUR_KEY" >> /etc/environment && source /etc/ environment echo export AOAI_ENDPOINT="YOUR_ENDPOINT" >> /etc/environment && ➥source /etc/environment NOTE In this book, we will use conda, an open source package manager, to manage our specific runtime versions and dependencies. Technically, using a package manager like conda is not mandatory, but it is extremely beneficial for isolating and troubleshooting problems and is highly recommended. We won’t delve into the specifics of installing conda in this context; for detailed, 62 CHAPTER 3 Working through an API: Generating text step-by-step instructions on how to install it, please refer to the official documentation at https://docs.conda.io/. First, let us create a new conda environment and install the required OpenAI Python library: $ conda create -n openai python=3.11.3 (base) $ conda activate openai (openai) $ conda install -c conda-forge openai Now that we have our dependencies installed, let’s connect to the Azure OpenAI endpoint and get details of the available models. 3.1.2 Listing models As we outlined earlier, each organization may have different models for use. We’ll start by understanding what models we have access to; we’ll use the APIs to help us set up the basic environment and get it running. Then, I’ll show you how to do this using the Azure OpenAI Python SDK and outline the differences when using the OpenAI API. As the next listing shows, we connect to the Azure OpenAI endpoint, get a list of all the models available, iterate over those, and print out the details of each model to the console. import os import json from openai import AzureOpenAI client = AzureOpenAI( azure_endpoint=os.getenv("AOAI_ENDPOINT"), api_version="2023-05-15", api_key=os.getenv("AOAI_KEY") ) # Call the models API to retrieve a list of available models models = client.models.list() # save to file with open('azure-oai-models.json', 'w') as file: models_dict = [model.__dict__ for model in models] json.dump(models_dict, file) # Print out the names of all the available models, and their capabilities for model in models: print("ID:", model.id) print("Current status:", model.lifecycle_status) print("Model capabilities:", model.capabilities) print("-------------------") Running this code will present us with a list of available models. The following listing shows an example of the models available; the exact list may be different for you. Listing 3.1 Listing Azure OpenAI models available Required for Azure OpenAI endpoints This is the environment variable pointing to the endpoint published via the Azure portal. Choose the API version we want to use from the multiple options. This is the environment variable with the API key. 633.1 Model categories { "id": "gpt-4-vision-preview", "created": null, "object": "model", "owned_by": null }, { "id": "dall-e-3", "created": null, "object": "model", "owned_by": null }, { "id": "gpt-35-turbo", "created": null, "object": "model", "owned_by": null }, … Each model is characterized by its distinct capabilities, suggesting the use cases for which it is tailored—specifically for chat completions, completions (which are regular text completions), embeddings, and fine-tuning. For example, a chat completion model would be the ideal selection in a situation where conversational engagement is required, like a chat-based interaction that requires significant dialogue exchange. Conversely, a completion model would be the most suitable for text generation. We can view the OpenAI base models with Azure AI Studio in figure 3.1. Figure 3.1 Base model listed Listing 3.2 Listing Azure OpenAI models’ output 70 CHAPTER 3 Working through an API: Generating text for choice in response.choices: print(choice.text) When we run this updated code, we get the response shown in listing 3.6. The property choices are an array, and we have three items, with the index starting at a base zero. Each has the generated text for us to use. Depending on the use case, this is helpful when picking multiple completions. 1. Pet Pampering Palace 2. Pet Grooming Haven 3. Perfect Pet Parlor 1. Pawsitive Pet Spa 2. Fur-Ever Friends Pet Salon 3. Purrfection Pet Care 1. Pampered Paws Professional Pet Care 2. Personalized Pet Pampering 3. Friendly Furrific Pet Care Another similar but more powerful parameter is the best_of parameter. Like the n parameter, it generates multiple completions, allowing the option to pick the best. The best_of is the completion with the highest log probability per token. We cannot stream results when using this option. However, it can be combined with the n parameters, with best_of needs greater than n. As shown in the following listing, if we set n to 5, we get five completions as expected; for brevity, we do not show all five of the completions here, but note that this call uses 184 tokens. { "choices": [ { … ], "created": 1689097645, "id": "cmpl-7bBkLk60mA8R9crAKXqTmTwzx2IEI", "model": "gpt-35-turbo", "object": "text_completion", "usage": { "completion_tokens": 152, "prompt_tokens": 32, "total_tokens": 184 } } Listing 3.6 Output showing multiple responses Listing 3.7 Output showing multiple responses 713.2 Completion API If we run a similar call using the best_of parameter, do not specify the n parameter: response = client.completions.create( model="gpt-35-turbo", prompt=prompt_startphrase, temperature=0.7, max_tokens=100, best_of=5, stop=None) When we run this code, we get only one completion, as shown in listing 3.8; however, we are using a similar number of tokens as earlier (171 versus 184). This is because the service generates five completions on the server side and returns the best one. The API uses the log probability per token to pick the best option. The higher the log probability, the more confident the model is about its prediction. { "choices": [ { "finish_reason": "stop", "index": 0, "logprobs": null, "text": "\n\n1. Pawsitively Professional Pet Salon\n ➥2. Friendly Furr Friends Pet Salon\n ➥3. Personalized Pampered Pets Salon", "content_filter_results"={...} } ], "created": 1689098048, "id": "cmpl-7bBqqpfuoV5nrgHrahuWGVAiM50Aj", "model": "gpt35", "object": "text_completion", "usage": { "completion_tokens": 139, "prompt_tokens": 32, "total_tokens": 171 } } The one parameter that influences many of the responses is the temperature setting. Let’s see how this changes the output. 3.2.4 Controlling randomness As discussed in the previous chapter, the temperature setting influences the randomness of the generated output. A lower temperature produces more repetitive and deterministic responses, while a higher temperature produces more innovative responses. Fundamentally, there isn’t a right setting—it all comes down to the use cases. Listing 3.8 Output generation with best_of five completions 72 CHAPTER 3 Working through an API: Generating text For enterprises, a more creative output would be when there is interest in diverse output and creating text for use cases such as content generation for marketing, stories, poems, lyrics, jokes, etc. These are things that usually require creativity. However, enterprises need more reliable and precise answers for use cases, such as document automation for invoice generation, proposals, code generation, etc. These settings are applicable per API call, so combining different temperature levels in the same workflow is possible. As demonstrated in previous examples, we recommend a temperature setting of 0.8 for creative responses. Conversely, a setting of 0.2 is suggested for more predictable responses. Using an example, let us examine how these settings alter the output and observe the variations between multiple calls. When the temperature was set to 0.8, we received the following responses from three consecutive calls. The output changes as expected, offering suggestions like those seen throughout this chapter. It is important to note that we do not need to make three separate API calls. We can set the n parameter to 3 in a single API call to generate multiple responses. Here is what our API call looks like: response = client.completions.create( model="gpt-35-turbo", prompt=prompt_startphrase, temperature=0.8, max_tokens=100, n=3, stop=None) The following listing shows the creative generation for the three responses. { "choices": [ { "finish_reason": "content_filter", "index": 0, "logprobs": null, "text": "", "content_filter_results"={...} }, { "finish_reason": "stop", "index": 1, "logprobs": null, "text": "\n\n1. Pawsitively Professional Pet Styling\n ➥2. Fur-Ever Friendly Pet Groomers \n ➥3. Tailored TLC Pet Care", "content_filter_results"={...} }, { "finish_reason": "stop", "index": 2, Listing 3.9 Completions output with the temperature at 0.8 First response: get blocked by the content filter Second of three responses Final response with very different generated text 733.2 Completion API "logprobs": null, "text": "\n\n1. Pawsitively Professional Pet Salon \n ➥2. Friendly Fur-ternity Pet Care \n ➥3. Personalized Pup Pampering Place", "content_filter_results"={...} } ], "created": 1689123394, "id": "cmpl-7bIRe6Ponn8y1198flJFfagq64r2E", "model": "gpt35", "object": "text_completion", "usage": { "completion_tokens": 96, "prompt_tokens": 32, "total_tokens": 128 } } Let’s change the setting to make this more deterministic and run it again. Note that the only change in the API call is temperature=0.2. The output is predictable and deterministic, with very similar text generated between the three responses. { "choices": [ { "finish_reason": "stop", "index": 0, "logprobs": null, "text": "\n\n1. Pawsitively Professional Pet Salon\n ➥2. Friendly Furr Salon\n ➥3. Personalized Pet Pampering", "content_filter_results"={...} }, { "finish_reason": "stop", "index": 1, "logprobs": null, "text": "\n\n1. Pawsitively Professional Pet Salon\n ➥2. Friendly Fur-Ever Pet Salon\n ➥3. Personalized Pet Pampering Salon", "content_filter_results"={...} }, { "finish_reason": "stop", "index": 2, "logprobs": null, "text": "\n\n1. Pampered Paws Pet Salon\n ➥2. Friendly Fur Salon\n ➥3. Professional Pet Pampering" } ], ... } Listing 3.10 Completions output with the temperature at 0.2 One of three responses Two of three responses; very similar generated text The final response with very similar generated text 74 CHAPTER 3 Working through an API: Generating text The temperature value goes up to 2, but it is not recommended to go that high, as the model starts hallucinating more and creating nonsensical text. If we want more creativity, we usually want it to be at 0.8 and, at most, 1.2. Let us see an example when the temperature is changed to 1.8. In this example, we did not even get the third generation, as we hit the token limit and stopped the generation. { "choices": [ { "finish_reason": "stop", "index": 0, "logprobs": null, "text": "\n\n1. ComfortGroom Pet Furnishing \n2. Pampered TreaBankant Carers \n3. Toptech Sunny Haven Promotion.", "content_filter_results"={...} }, { "finish_reason": "stop", "index": 1, "logprobs": null, "text": "\n\n1: Naturalistov ClearlywowGroomingz ➥Pet Luxusia \n2: VipalMinderers Pet ➥Starencatines grooming \n3: Brisasia ➥Crownsnus Take Care Buddsroshesipalising", "content_filter_results"={...} }, { "finish_reason": "length", "index": 2, "logprobs": null, "text": "\n\n1. TrustowStar Pet Salon\n ➥2. Hartipad TailTagz Grooming & Styles\n ➥3. LittleLoft Millonista Cosmania DipSavez ➥Hubopolis ShineBright Princessly ➥Prosnoiffarianistics Kensoph Cowlosophy ➥Expressionala Navixfordti Mundulante Effority ➥DivineSponn BordloveDV EnityzBFA Prestageinato ➥SuperGold Cloutoilyna Critinarillies ➥Prochromomumphance Toud", ➥"content_filter_results"={...} } ], ... } 3.2.5 Controlling randomness using top_p An alternative to the temperature parameter for managing randomness is the top_p parameter. It has the same affect on the generation as the temperature parameter, but it uses a different technique called nucleus sampling. Essentially, nucleus sampling Listing 3.11 Completions output with the temperature at 1.8 One of three responses with names that aren’t very clear Second and third of three responses, with nonsensical names 753.3 Advanced completion API options allows only the tokens with a probability equal to or less than the value of top_p to be considered as part of the generation. Nucleus sampling creates texts by picking words from a small group of the most likely ones with the highest cumulative probability. The top_p value decides how small this group is based on the total chance for the words to appear in it. The group size can change depending on the next word’s chance. Nucleus sampling can help avoid repetition and generate more varied and clearer texts than other methods. For example, if we have the top_p value set to 0.9, only the tokens that make up 90% of the probability distribution will be sampled for the generation of text. This allows us to avoid the last 10%, which are often quite random and diverse and end up as nonsensical hallucinations. A lower value of top_p makes the model more consistent and less creative as it chooses fewer tokens to generate. Conversely, a higher value makes the generation more creative and diverse, as it has a larger set of tokens to operate. The larger value also makes it prone to more errors and randomness. The exact value of top_p depends on the use case; in most cases, the ideal value for top_p ranges between 0.7 and 0.95. We should change either the temperature attribute or top_p, but not both. Table 3.5 outlines the relationship between the two. Let us look at some of the advanced API options for specific scenarios. 3.3 Advanced completion API options Now that we have examined the basic constructs of the completion API and understand how they work, we need to consider more advanced aspects of the completion API. Many of these might not seem as complex, but they add many more responsibilities to the system architecture, complicating overall implementation. 3.3.1 Streaming completions The completions API allows streaming responses, offering immediate access to information as soon as it is ready rather than waiting for a full response. For enterprises, streaming can be important in some cases where real-time content generation with Table 3.5 Relationship between temperature and top_p Temperature top_p Effect Low Low Generates predictable text that closely follows common language patterns Low High Generates predictable text, but with occasional less common words or phrases High Low Generates text that is often coherent but with creative and unexpected word usage High High Generates highly diverse and unpredictable text with various word choices and ideas; has very creative and diverse output, but may contain many errors 76 CHAPTER 3 Working through an API: Generating text lower latency is key. This feature can enhance user experiences by processing incoming responses promptly. To enable streaming from the API’s standpoint, modify the stream parameter to true. By default, this optional parameter is set to false. Streaming employs server-sent events (SSE), which require a client-side implementation. SSE is a standard protocol allowing servers to continue transmitting data to clients after establishing the initial connection. It is a long-term, one-way connection from server to client. SSE offers advantages such as low latency, reduced bandwidth consumption, and an uncomplicated configuration setup. Listing 3.12 demonstrates how our example can be adjusted to utilize streaming. Although the API modification is straightforward, the description and requested multiple generations were adjusted (using the n property). This allows us to generate more text artificially, making it easier to observe the streaming generation. import os import sys from openai import AzureOpenAI client = AzureOpenAI( azure_endpoint=os.getenv("AOAI_ENDPOINT"), api_version="2024-05-01-preview", api_key=os.getenv("AOAI_KEY")) prompt_startphrase = "Suggest three names and a tagline ➥which is at least 3 sentences for a new pet salon business. ➥The generated name ideas should evoke positive emotions and the ➥followingkey features: Professional, friendly, Personalized Service." for response in client.completions.create( model="gpt-35-turbo", prompt=prompt_startphrase, temperature=0.8, max_tokens=500, stream=True, stop=None): for choice in response.choices: sys.stdout.write(str(choice.text)+"\n") sys.stdout.flush() When managing a streaming call, we must pay extra attention to the finish_reason property. As messages are streamed, each appears as a standard completion, with the text representing the newly generated token. In these instances, the finish_reason remains null. However, the final message differs; its finish_reason could be either stop or length, depending on what triggered it. Listing 3.12 Streaming completion Tweaked the prompt slightly to add descriptions We need to handle the streaming response on the client side. Enables streaming We need to loop through the array and handle multiple generations. 773.3 Advanced completion API options ... { "finish_reason": null, "index": 0, "logprobs": null, "text": " Pet" } { "finish_reason": null, "index": 0, "logprobs": null, "text": " Pam" } { "finish_reason": null, "index": 0, "logprobs": null, "text": "pering" } { "finish_reason": "stop", "index": 0, "logprobs": null, "text": "" } 3.3.2 Influencing token probabilities: logit_bias The logit_bias parameter is one way we can influence output completion. In the API, this parameter allows us to manipulate the probability of certain tokens, which can be words or phrases, that the model generates in its responses. It is called logit_ bias because it directly affects the log odds, or logits, that the model calculates for each potential token during the generation process. The bias values are added to these log-odds before converting them to probabilities, altering the final distribution of tokens the model can pick from. The importance of this feature lies in its ability to steer the model’s output. Say we are creating a chatbot and want it to avoid certain words or phrases. We can use logit_bias to decrease the likelihood of those tokens being chosen by the model. In contrast, if there are certain words or phrases we want the model to favor, we could use logit_bias to increase their likelihood. The range of this parameter is from –100 to 100, and it operates on tokens for the word. Setting a token to –100 effectively bans it from the generation, whereas setting it to 100 makes it exclusive. To use logit_bias, we provide a dictionary where the keys are the tokens, and the val-ues are the biases that need to be applied to those tokens. To get the token, we use the tiktoken library. Once you have the appropriate token, you can assign a positive bias to make it more likely to appear or a negative bias to make it less likely, as shown in figure 3.2. The blocks show the degree of probability that different tokens can be at different Listing 3.13 Streaming finish reason 78 CHAPTER 3 Working through an API: Generating text probabilities of banning or exclusive generation. Smaller changes to the tokens’ value increase or decrease the probability of these tokens in the generated output. Figure 3.2 The logit_bias parameter Let’s use an example to see how we can make this work. For our pet salon name, we do not want to use the words “purr,” “purrs,” or “meow.” The first thing we want to do is create the tokens for these words. We also want to add words with a preceding space and capitalize them as spaces. Capital letters are all different tokens. So “Meow” and “Meow” (with a space) and “meow” (again with a space) might read the same to us, but when it comes to tokens, these words are all different. The output shows us the tokens for the corresponding word: 'Purr Purrs Meow Purr purr purrs meow:[30026, 81, 9330, ➥3808, 42114, 9330, 81, 1308, 81, 1308, 3808, 502, 322]' Now that we have the tokens, we can add them to the completion call. Note that we assign each token a bias of –100, steering the model away from these words. import os from openai import AzureOpenAI client = AzureOpenAI( azure_endpoint=os.getenv("AOAI_ENDPOINT"), api_version="2024-05-01-preview", api_key=os.getenv("AOAI_KEY")) GPT_MODEL = "gpt-35-turbo" prompt_startphrase = "Suggest three names for a new pet salon ➥business. The generated name ideas should evoke positive ➥emotions and the following key features: Professional, ➥friendly, Personalized Service." response = client.completions( model=GPT_MODEL, Listing 3.14 logit_bias implementation Tokens Ban Exclusive Output logit_bias -100 100 Lower probability Higher probability 793.3 Advanced completion API options prompt=prompt_startphrase, temperature=0.8, max_tokens=100, logit_bias={ 30026:-100, 81:-100, 9330:-100, 808:-100, 42114:-100, 1308:-100, 3808:-100, 502:-100, 322:-100 } ) responsetext =response.choices[0].text print("Prompt:" + prompt_startphrase + "\nResponse:" + responsetext) We do not have any words we want to avoid when we run this code. { "choices": [ { "finish_reason": "stop", "index": 0, "logprobs": null, "text": "\n\n1. Paw Prints Pet Pampering\n2. Furry Friends Fussing\n3. Posh Pet Pooches" } ], ... } We can do the opposite and positively bias tokens too. Say we want to overemphasize and steer the model toward the word “Furry.” We can use the tiktoken library we saw earlier and find that the tokens for “Furry” are [37, 16682]. We can update the previous API call with this and, in this case, a positive bias of 5. GPT_MODEL = "gpt-35-turbo" response = client.completions.create( model=GPT_MODEL, prompt=prompt_startphrase, temperature=0.8, max_tokens=100, logit_bias={ 30026:-100, 81:-100, Listing 3.15 Output of logit_bias generation Listing 3.16 logit_bias: Positive implementation Dictionary containing the tokens and the corresponding bias values to steer the model on these specific tokens 86 CHAPTER 3 Working through an API: Generating text NOTE The following parameters are unavailable with the new GPT-35 Turbo and GPT-4 models: logprobs, best_of, and echo. Trying to set any of these parameters will throw an exception. The output of the previous example is shown in the next listing. The user started with “Hello, World!”, and the system responded, asking how to help us with the assistant message. The question about dog details is the next dialogue turn. { "id": "chatcmpl-7c9HOaRHmaut94oZacjdUqxOL96H3", "object": "chat.completion", "created": 1689326490, "model": "gpt-35-turbo", "choices": [ { "index": 0, "finish_reason": "stop", "message": { "role": "assistant", "content": "Dogs are often considered man's best ➥friend for good reason. They provide ➥companionship, protection, and loyalty to ➥their owners. Here are some reasons why dogs are ➥good for humans:\n\n1. They reduce stress: Dogs ➥have been shown to reduce stress and anxiety in ➥their owners. Petting a dog can release endorphins, ➥which are feel-good hormones that can help to ➥reduce stress.\n\n2. They provide companionship..." } } ], "usage": { "completion_tokens": 238, "prompt_tokens": 60, "total_tokens": 298 } } 3.4.1 System role The system role (some also call it the system message) is included at the beginning of the message array. This message provides the initial instructions for the model, and we can provide various pieces of information in the system role, including A brief description of the assistant Personality traits of the assistant Rules and instructions you want the assistant to follow Additional information needed for the model (e.g., relevant questions from an FAQ) Listing 3.20 Chat completion API output 873.4 Chat completion API We customize the system role and include basic instructions for the use case. From an API perspective, even though the system role is optional, it is highly recommended that you make this intentional to get the best results. For example, if we expand on the previous example of chatting for pets and pet salons, we can instruct the model to only reply in rhyme. GPT_MODEL = "gpt-35-turbo" response = client.chat.completions.create( engine=GPT_MODEL, messages = [ {"role": "system", "content": "You are a helpful AI ➥assistant that provides truthful information. ➥You answer all questions in rhyme."}, {"role": "user", "content": "Hi"}, {"role": "assistant", "content": "Hello there, how can ➥I assist?\nAsk me a question, don't resist!"}, {"role": "user", "content": "who are the founders of Microsoft?"}, {"role": "assistant", "content": "Bill Gates and Paul Allen, ➥it's true,\nAre the founders of Microsoft, ➥through and through."}, {"role": "user", "content": "what is a good name for a pet salon?"}, {"role": "assistant","content": "For a pet salon [ ➥that's simply divine,\nHere's a name that's [ ➥sure to shine:\n\"Paws and Pamper\" is what I ➥propose,\nA name that's catchy and easy to compose."}], temperature=0.7, max_tokens=800, top_p=0.95, frequency_penalty=0, presence_penalty=0, stop=None) In the example, we can have a conversation as expected, which can vary topics in turns, but all the answers rhyme. When we want to give the model additional data as context for the conversation, this is called grounding the data. If there is a small amount of data, this can be part of the system role, as shown in the next listing. However, if there is a large amount of data, we should use embeddings and retrieve the most relevant information using a semantic search (e.g., Azure cognitive search). {"role": "system", "content": Assistant is an intelligent chatbot designed to help users answer technical questions about Azure OpenAI service. Only answer questions using the context below. Don’t make up the answer. If you are unsure of an answer, say 'I don't know'. Listing 3.21 Chat completion system message example Listing 3.22 Grounding system message example Instructs to answer in rhyme 88 CHAPTER 3 Working through an API: Generating text Context: - Azure OpenAI Service provides REST API access to OpenAI's powerful language models including the GPT-3, Codex and Embeddings model series. - Azure OpenAI Service gives customers advanced language AI with OpenAI GPT3, Codex, and DALL-E models with the security and enterprise promise of Azure. ..." }, {"role": "user", "content": "What is Azure OpenAI Service?"} 3.4.2 Finish reason Every chat completion API response has a finish reason encoded in the finish_ reason field. Tracking is important in this case, as it helps us understand why the API returned the response it did. This can be useful for debugging and improving the application. For example, if you receive an incomplete response due to the length finish reason, you may want to adjust the max_tokens parameter to generate more complete responses. The possible values for finish_reason are stop—The API finished generating and either returned a complete message or a message terminated by one of the stop sequences provided using the stop parameter. length—The API stopped the model output due to the max_tokens parameter or token limit. function_call—The model decided to call a function. content_filter—Some of the completion was filtered due to harmful content. 3.4.3 Chat completion API for nonchat scenarios OpenAI’s chat completion can be used for nonchat scenarios. The API is quite similar and designed to be a flexible tool that can be adapted to various use cases, not just conversations. In most cases, the recommended path uses the chat completion API as if it were the completion API. The main reason is that the newer models (Chat 3.5Turbo and GPT-4) are much more efficient, cheaper, and powerful than the earlier models. The completion use cases we have seen, such as analyzing and generating text and answering questions from a knowledge base, would all still work with the chat completion API. Implementing the chat completion API nonchat scenarios usually involves structuring the conversation with a series of messages and a system message to set the assistant’s behavior. For example, as shown in the following listing, the system message sets the role of the assistant, and the user message provides the task. GPT_MODEL = "gpt-35-turbo" response = client.chat.completions.create( model=GPT_MODEL, Listing 3.23 Chat completion as a completion API example 893.4 Chat completion API messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Translate the following ➥English text to Spanish: 'Hello, how are you?'"} ] ) We can also use a series of user messages to provide more context or accomplish more complex tasks, as shown in the next listing. In this example, the first user message sets up the task, and the second user message provides more specific details. The assistant generates a response that attempts to complete the task in the user messages. GPT_MODEL = "gpt-35-turbo" response = client.chat.completions.create( model=GPT_MODEL, messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "I need to write a Python function."}, {"role": "user", "content": "This function should take two ➥numbers as input and return their sum."} ] ) 3.4.4 Managing conversation Our examples keep running, but the conversation will hit the model’s token limit as it continues. With each turn of the conversation (i.e., the question asked and the answer received), the list of messages grows. As a reminder, the token limit for GPT-35 Turbo is 4K tokens, and for GPT-4 and GPT-4 32K, it is 8K and 32K, respectively; these include the total count from the message list sent and the model response. We get an exception if the total count exceeds the relevant model limit. No out-of-the-box option can track this token count for us and ensure it falls within the token limit. As part of the enterprise app design, we need to track the token count and only send a prompt that falls within the limit. Many enterprises are in the process of implementing an enterprise version of ChatGPT using the chat API. Here are some of the best practices that can help enterprises manage these conversations. Remember, the best way to get your desired output involves iterative testing and refining your instructions: Setting the behavior with system message—You should use the system message at the start of the conversation to guide the model’s behavior and for enterprises to tune to reflect their brand or IP. Providing explicit instructions—If the model is not generating your desired output, make your instructions more explicit. Think about it at the same level as if you were telling a toddler what not to do. Listing 3.24 Chat completion as a completion API example 90 CHAPTER 3 Working through an API: Generating text Breaking down complex tasks—If you have a complex task, break it down into several simpler tasks, and send them as separate user messages. You often need to show, not explain it. This is called Chain of Thought (CoT), and it will be covered in more detail in chapter 6. Experimentation—Feel free to experiment with the parameters to get the desired output. A higher temperature value (e.g., 0.8) makes the generation more random, while a lower value (e.g., 0.2) makes it more deterministic. You can also use the maximum token value to limit response length. Managing tokens—Be aware of the total number of tokens in a conversation, as input and output tokens count toward the total. You must truncate, omit, or shorten your text if a conversation has too many tokens to fit within the model's maximum limit. Handling sensitive content—If you’re dealing with potentially unsafe content, you should look at Azure OpenAI’s Responsible AI guidelines (https://mng.bz/ pxVK). However, if you are using OpenAI’s API, then OpenAI’s moderation guide is helpful (https://mng.bz/OmEw) for adding a moderation layer to the outputs of the chat API. TRACKING TOKENS As outlined earlier, keeping track of tokens when using the conversational API is key. Not only will the experience suffer if we go over the total token size, but the total number of tokens in an API also has a direct effect on latency and on how long the call takes. Finally, the more tokens we use, the more we pay. Here are some ways you can manage tokens: Count tokens. Use the tiktoken library, which allows us to count how many tokens are in a string without making an API call. Limit response length. When making an API call, use the max_tokens property to limit the length of the model’s responses. Truncate long conversations. If a conversation has too many tokens to fit within the model’s maximum limit, we must truncate, omit, or shorten our text. Limit the number of turns. Limiting the number of turns in the conversation is a good way to truncate or shorten the text. This also helps steer the model better when the conversation gets longer and tends to start hallucinating. Check the usage field in the API response. After making an API call, we can check the usage field in the API response to see the total number of tokens used. This is ongoing and includes both input and output tokens. It is a good way to keep track of tokens and show them to the user via some UX. Reduce temperature. Reducing the temperature parameter can make the model's outputs more focused and concise, which can help reduce the number of tokens used in the response. Say we want to build a chat application for our pet salon and allow customers to ask us questions about pets, grooming, and their needs. We can build a console chat 913.4 Chat completion API application, as shown in listing 3.25. It also shows us a possible way to track and manage tokens. In this example, we have a function num_tokens_from_messages which, as the name suggests, is used to calculate the number of tokens in a conversation. As the conversation grows turn by turn, we calculate the number of tokens used, and once it reaches the model limit, the old messages are removed from the conversation. Note that we start at index 1. This ensures we always preserve the system message at index 0 and only remove user/assistant messages. import os from openai import AzureOpenAI import tiktoken client = AzureOpenAI( azure_endpoint=os.getenv("AOAI_ENDPOINT"), api_version=”2024-05-01-preview”, api_key=os.getenv(“AOAI_KEY”)) GPT_MODEL = "gpt-35-turbo" system_message = {"role": "system", "content": "You are ➥a helpful assistant max_response_tokens = 250 token_limit = 4096 conversation = [] conversation.append(system_message) def num_tokens_from_messages(messages): encoding= tiktoken.get_encoding("cl100k_base") num_tokens = 0 for message in messages: num_tokens += 4 for key, value in message.items(): num_tokens += len(encoding.encode(value)) if key == "name": num_tokens += -1 num_tokens += 2 print("I am a helpful assistant. I can talk about pets and salons.") while True: user_input = input("") conversation.append({"role": "user", "content": user_input}) conv_history_tokens = num_tokens_from_messages(conversation) while conv_history_tokens + max_response_tokens >= token_limit: del conversation[1] conv_history_tokens = num_tokens_from_messages(conversation) response = client.chat.completions.create( model=GPT_MODEL, Listing 3.25 ConsoleChatApp: Token management Sets up the OpenAI environment and configuration details Sets up the system message for the chat Function to count the total tokens from all the messages in the conversation Uses the tiktoken library to count tokens Loops through the messages Captures the user input When the total tokens exceed the token limit, we remove the second token. The first token is the system token, which we always want. Chat completion API call 92 CHAPTER 3 Working through an API: Generating text messages=conversation, temperature=0.8, max_tokens=max_response_tokens) conversation.append({"role": "assistant", "content": ➥response.choices[0].message.content}) print("\n" + response.choices[0].message.content) print("(Tokens used: " + str(response.usage.total_tokens) + ")") CHAT COMPLETION VS. COMPLETION API Both chat completion and completion APIs are designed to generate human-like text and are used in different contexts. The completion API is designed for single-turn tasks, providing completion to a prompt provided by the user. It is most suited for tasks where only a single response is required. In contrast, the chat completion API is designed for multiturn conversations, maintaining the context of the conversation over multiple exchanges. This makes it more suitable for interactive applications such as chatbots. The chat completion API is a new dedicated API for interacting with the GPT-35-Turbo and GPT-4 models and is the preferred method. The chat completion API is geared more toward chatbots, and using the different roles (system, user, and assistant), we can get the memory of previous messages and organize few-shot examples. 3.4.5 Best practices for managing tokens For LLMs, tokens are the new currency. As most enterprises go beyond kicking tires to business-critical use cases, managing tokens would become a priority for computations, cost, and overall experience. From an enterprise application perspective, here are some of the considerations for managing tokens: Concise prompts—Where possible, using concise prompts and limiting the maximum number of tokens will reduce the token’s usage, making it more costeffective. Stop sequences—Use stop sequences to stop the generations to avoid generating unnecessary tokens. Counting tokens—We can count tokens using the tiktoken library as outlined earlier and avoid making the API calls do the same. Smaller models—Generally speaking, in computing, bigger and newer hardware and software are considered faster, cheaper, and better; however, this isn’t necessarily the case for LLMs. Where possible, consider using smaller models such as GPT-3.5 Turbo first, and when they might not be a good fit, consider going to the next one. Smaller models are less compute intensive and, hence, are more economical. Use caching—For prompts that are either quite static or frequently repeated, implementing a caching strategy would help save tokens and avoid making API calls repeatedly. For more complex scenarios, look to cache the embeddings 933.4 Chat completion API using a vector search and store, such as Azure Cognitive Search, Pinecone, etc. The last chapter covered an introduction to embeddings, and we will get more details on embeddings and searching later in chapters 7 and 8 when we cover RAG and chatting with your data. 3.4.6 Additional LLM providers Additional vendors also now have LLMs to use for enterprises. These are either available via APIs or, in some cases, as model weights that enterprises can self-host. Table 3.7 outlines some of the more famous ones available at the time of publication. Please note that some restrictions are in place from a commercial-licensing perspective. Interestingly, all these vendors follow a similar approach to the concepts and APIs established by OpenAI. For example, as outlined by their documents, the PaLM model from Google’s completion API equivalent is presented in the next listing. google.generativeai.generate_text(*, model: Optional[model_types.ModelNameOptions] = 'models/text-bison-001', prompt: str, temperature: Optional[float] = None, Table 3.7 Other LLM providers Models Descriptions Llama 2 Meta released Llama 2, an open source LLM, which comes in three sizes (7 billion, 13 billion, and 70 billion parameters) and is free for research and commercial purposes. Companies can access this through cloud options such as Azure AI’s model catalog, Hugging Face, or AWS. Enterprises that want to host it using their own compute and GPUs can request access from Meta via https://ai.meta.com/llama/. PaLM PaLM is a 13 billion-parameter model from Google that is part of their generative AI for developer products. The model can perform text summarization, dialogue generation, and natural language inference tasks. At the time of publication, there was a waitlist for an API key; details are available at https://developers.generativeai.google/. BLOOM Bloom is a 223-billion parameter, open source multilingual model that can understand and generate text in over 100 languages by collaborating with over 1,000 researchers across more than 250 institutions. It is available via Hugging Face for deployment. More details are available at https://huggingface.co/bigscience/bloom. Claude Claude is a 12-billion parameter developed by Anthropic. It is accessible through a playground interface and API in its developer console for development and evaluation purposes only. At publication, for production use, enterprises must contact Claude for commercial discussions. More details can be found at https://mng.bz/YVqz. Gemini Google recently released a new LLM called Gemini, a successor to PaLM 2 and optimized for different sizes: ultra, pro, and nano. It is designed to be more powerful than its predecessor and can be used to generate new content. Google claims it to be their most capable AI model yet. More details can be found at https://mng.bz/GNxD. Listing 3.26 PaLM-generated text API signature 94 CHAPTER 3 Working through an API: Generating text max_output_tokens: Optional[int] = None, top_p: Optional[float] = None, top_k: Optional[float] = None, stop_sequences: Union[str, Iterable[str]] = None, ) -> text_types.Completion While these options exist, and some are from reputable and leading technology companies, for most enterprises, Azure OpenAI and OpenAI are the most mature, with the most enterprise controls and support needed. The next chapter will deal with images, and we will learn how to move from text to images and generate in that modality. Summary GenAI models are classified into various categories, depending on the type. Each model has additional capabilities and characteristics. Choosing the right model for the use case at hand is important. And unlike computer science, in our case, the biggest model isn’t necessarily better. The completion API is a sophisticated tool that generates text, which can be used to complete prompts provided by the user and forms the backbone of the text generation paradigm. The completion API is relatively easy to use with only a few key parameters, such as the prompt, number of tokens to generate, temperature parameter that helps steer the model, and number of completions to generate. The API exposes many advanced options for steering models and controlling randomness and generated text, such as logit_bias, presence penalty, and frequency penalty. All these work in tandem and help generate better output. When using Azure OpenAI, the content safety filter can help filter specific categories to identify and act on potentially harmful content as part of both the input prompts and generated completions. The chat completion API builds on the completion API, going from one set of instructions and APIs to a dialogue with the user in a turn-by-turn interaction. The chat completion consists of multiple systems, user, and assistance roles. The conversation starts with a system message that sets the assistant's behavior, followed by alternating user and assistant messages as the conversation proceeds turn by turn. The system role is included at the beginning of the message array. It provides the initial instructions for the model, including personality traits, instructions and rules for the assistant to follow, and additional information we want to provide as context for the model; this additional information is called grounding the data. Each completion and chat completion API response has a finish reason, which helps us understand why the API returned the response it did. This can be useful for debugging and improving the application. 95Summary The language learning models all have a finite context window and are quite expensive. Managing tokens becomes important for us to be able to run things at a reasonable cost and within the API allowance. This also helps us manage tokens in conversations for improved user experience and cost-effectiveness. In addition to Azure OpenAI and OpenAI, there are other LLM providers, such as Meta’s Llama 2, Google’s Gemini and PaLM, Bloom by BigScience, and Anthropic’s Claude. Their offerings are similar and follow the completions and chat completions paradigm, including similar APIs. 102 CHAPTER 4 From pixels to pictures: Generating images Figure 4.5 GAN model architecture GANs offer many similar use cases, such as VAEs, but they are specifically good for Image generation—Creating realistic images from noise, with specific applications in entertaining, design, and art, allows generating high-quality images. Style transfer—Enabling artistic styles to transpose from one image to another; this is the same as in VAEs. Super resolutions—GANs can help enhance resolution, making images more detailed and clearer. This is very helpful in some industries, such as medical and space imaging. Data augmentation—Similar to VAEs for creating synthetic data, GANs help create training data either for edge cases or where there is not enough data or data diversity. GANs can produce high-quality images that are indistinguishable from real ones. Still, they have drawbacks, such as mode collapse (i.e., the model repeatedly produces the same output), instability, and difficulty controlling the output. They also raise ethical concerns, as they can be quite easily used to create deepfakes that could lead to privacy invasion, potential misinformation, and misrepresentation. Finally, as with many other AI models, GANs can inadvertently perpetuate biases present in the training data in the generated output. 4.1.3 Vision transformer models Transformers are another model architecture that can create images. We saw the same architecture earlier in the context of natural language processing (NLP) tasks. Transformers can also operate on vision-related tasks and are called vision transformers (ViT) [2]. Transformers are neural networks that use attention mechanisms to process sequential data, such as text or speech, and they can be used to generate image Latent space G generator D discriminator Noise Learn data distribution Generate fake samples Fine-tuning training Learn difference between fake samples and real samples Real samples Is D correct? 1034.1 Vision models prompts. They are also very effective for specific tasks such as image recognition and have outperformed previous leading model architectures. A ViT model’s architecture is similar to that of NLP, albeit with some differences— it has a larger number of self-attention layers and a global attention mechanism allowing the model to attend to all parts of the image simultaneously. Transformers calculate how much each input token is related to every other input token. This is called attention. The more tokens there are, the more attention calculations are needed. The number of attention calculations grows as the square of the number of tokens, that is, is quadratically. For images, however, the basic unit of analysis is a pixel and not a token. The relationships for every pixel pair in a typical image are computationally prohibitive. Instead, ViT computes relationships among pixels in various small sections of the image (typically in 16 × 16-sized pixels), which helps reduce the computational cost. These 16 × 16-sized sections, along with their positional embeddings, are placed in a linear sequence and are the input to the transformer. As shown in figure 4.6, a ViT model consists of three main sections: the left, middle, and right. The left section shows the input classes, such as Class, Bird, Ball, Car, and so forth. These are the possible labels that the model can assign to an image. The middle section shows the linear projection of flattened patches, which transform the input image into a sequence of vectors that can be fed to the transformer encoder. The final section is the transformer encoder. This comprises several multi-head attention and normalization layers and is used to learn the relationships between different image parts. Figure 4.6 Vision transformer (ViT) architecture [2] Vision transformer (ViT) Transformer encoder MLP head Class Bird Ball Car ... Transformer encoder Linear projection of flattened patches Multihead attention MLP Norm Norm Embedded patches Patch + position embedding * Extra learnable [class] embedding 104 CHAPTER 4 From pixels to pictures: Generating images ViTs are used for various image use cases, such as segmentation, classification, and detection, and they are often more accurate than previous techniques. They also support fine-tuning, which can be used in a few-shot manner with smaller datasets, making them quite useful for enterprise use cases where we might not have much data. The ViT model aims to produce a final vector representation for the class token, which contains information about the whole image. ViTs also have challenges such as high computational costs, data scarcity, and ethical issues. They are computationally complex both from a training and inference perspective and have low interpretability—both active research areas. Multimodal models, with ViTs such as GPT-4, hold much promise and unlock new enterprise possibilities. 4.1.4 Diffusion models Diffusion models are generative machine learning models that can create realistic data from random noise, such as images or audio. Their goal is to learn the latent structure of a dataset by modeling how data points diffuse through that latent space. The model is trained by slowly adding noise to an image and learning to reverse this by removing noise from the input until it resembles the desired output. For example, a diffusion model can generate an image of a panda by starting with a random image and then slowly removing noise until it looks like a panda. Vision diffusion models typically consist of two parts: a forward and a reverse diffusion process. The forward diffusion process is responsible for gradually adding noise to the latent representation of an image, which corrupts that latent space. The reverse diffusion process is just the opposite—it is responsible for reconstructing the original image from the corrupted latent representation. The forward diffusion process is typically implemented as a Markov chain (i.e., a system with no memory of its past, and the probability of the next step depends on the current state). This means the corrupted latent representation at each step depends only on the previous step’s latent representation, which makes the forward diffusion process efficient and easy to train. The reverse diffusion process is typically implemented as a neural network, meaning the neural network learns to reverse the forward diffusion process by predicting the original latent representation from the corrupted one. This reverse diffusion process is slow, as it is a step-by-step repetition. Some of the advantages that diffusion models have are the following: They can produce high-quality images that match or beat GAN-generated images, especially for complex scenes, but they take much longer to generate. They do not suffer from mode collapse, a common problem for GANs. Mode collapse occurs when the generator produces only a limited variety of outputs, ignoring some modes of data distribution. Diffusion models can capture the full diversity of the data distribution by using a Markov chain process that adds noise to the input data. Diffusion models can be combined with others, such as natural language models, to create text-guided generation systems. 1054.1 Vision models Stable Diffusion is one of the most popular diffusion-based models for image generation. Its architecture consists of three main parts (see figure 4.7): The text encoder, which converts the user’s prompt into a vector representation. A denoising autoencoder (called UNet), which is used to reconstruct an image from the latency space, and a scheduler algorithm, which helps reconstruct the original image. We call it the image information creator. The UNet is a denoising autoencoder because it learns to remove noise from the input image and produce a clean output image. It is a neural network that has an encoder–decoder structure. The encoder part reduces the resolution of an input image and extracts its features. On the other hand, the decoder part increases the resolution and reconstructs the output image. A variational autoencoder (VAE), which creates an image as close as possible to a normal distribution. Figure 4.7 Stable Diffusion logical architecture Text conditioned latent UNet Text encoder (CLIP) Image decoder (variational autoencoder) Repeat Nsteps 64 x 64 Conditional latents 64 x 64 Latents 77 x 768 Text embeddings Prompt: “A panda riding a wave” Latent seed Random image 512 x 512 Generated image Image information creator Schedule algorithm “reconstruct” 1 2 3 106 CHAPTER 4 From pixels to pictures: Generating images The choice between these models depends on the specific application, the availability of computing resources, training data, and nonfunctional requirements such as image quality, speed, and so forth. Table 4.1 lists some of the more common generative AI vision systems that can create images from text. Many of the AI vision models listed in table 4.1 are available only to those who were invited to test them. This is still a new space, and most providers are going slowly, learning with a handful of customers before rolling these out. Creating and manipulating images with generative AI is an exciting and challenging research area with many potential applications and implications. However, it raises ethical and social questions about the generated content’s ownership, authenticity, and effects. Therefore, it is important to use generative AI responsibly and ethically and to consider its benefits and risks to society. 4.1.5 Multimodal models A multimodal model can handle different types of input data. “Modal” refers to the mode or type of data, and “multimodal” refers to multiple data types. These types include text, images, audio, video, and more. For example, GPT-4 has a multimodal Table 4.1 Most common AI vision tools AI vision tool Description Imagen Imagen is Google’s text-to-image diffusion model, which can generate realistic images from text descriptions. It is available in limited preview and has been shown to generate images indistinguishable from real photographs. DALL-E OpenAI developed a transformer language model to create diverse, original, realistic, and creative images and art from a prompt. It can edit images based on the context, such as adding, deleting, or changing specific parts. It has generated various images, from everyday objects to surrealistic art, from simple text prompts. DALL-E 3 is an improved version that can generate more realistic and accurate images with 4x greater resolution. Midjourney AI-based art generator that uses deep learning and neural networks to create artwork based on prompts and other images and videos. This is accessible only via a Discord server, and the results can be tailored to any aesthetics, from abstract to realistic, thus offering endless possibilities for creative expression. Adobe Firefly Adobe Firefly is a family of creative, generative AI diffusion models designed to help designers and creative professionals create images and text effects and edit and recolor. It is easy to use with Adobe’s other tools, such as Photoshop and Illustrator. Adobe has both text-to-image models and generative fill models. Stable Diffusion Popular models include versions of Stable Diffusion XL and v1.6, an imagegenerating model that uses diffusion models to create high-quality images using prompts with next-level photorealism capabilities. It can also generate novel images from text descriptions. The more recent v3 family of models comes in large and medium with 8B and 2B parameters, respectively. 1074.1 Vision models model variant that takes both an image and an associated prompt to make predictions or inferences. Bing Chat recently enabled this multimodal feature, allowing us to use images and text in the prompt. For example, as shown in figure 4.8, we give the model two things: an image and a prompt related to the image. In this case, we show some produce and ask the model what we can cook with it. Figure 4.8 Multimodal example using both an image and a prompt In this case, the model must understand the image and the different parts (i.e., ingredients in our example) and correlate to the prompt to generate an answer. We see the response in the shaded text, showing we can make guacamole, salsa, avocado toast, and so forth. 108 CHAPTER 4 From pixels to pictures: Generating images Multimodal models often use different AI techniques. While they can use different combinations of model architecture, in our example, GPT-4 combines different transformer blocks (figure 4.9). Figure 4.9 Multimodal model design NOTE When showing transformer blocks, as in figure 4.9, the convention is to use Nx, referring to the transformer block repeating multiple times; in other words, it is stacked x number of times. In our multimodal example, this is the case for all three transformer blocks: the image on the left (Lx), the text on the right (Rx), and the combining layer (Nx). Multi self-attention Feed forward Combining layer Visual embeddings Transformer block Multi self-attention Feed forward Text embeddings Transformer block Lx Nx Rx “What can I make with this?” Text input Image input + + + + Multi self-attention Feed forward Transformer block + + Layer norm Layer norm Layer norm Layer norm Layer norm Layer norm 1094.2 Image generation with Stable Diffusion Multimodal models are particularly useful in complex real-world applications where data comes in various forms. For example: Web—Analyzes text and images for content moderation and sentiment analysis eCommerce—Recommends products using both photos and text descriptions Healthcare—Uses text data (patient medical history) and medical imaging (image data) for diagnosis Self-driving—Integrates sensor data (radar and lidar) with visual data (cameras) for situational awareness and decision-making Now that we have seen some models, their output, and a general sense of how vision AI models work, let us generate images with Stable Diffusion. 4.2 Image generation with Stable Diffusion Stability AI, the company behind the Stable Diffusion, has advanced diffusion-based models with SDXL as their latest and most powerful model thus far. They offer multiple options for us to use: Self-host—The model and associated weights have been published and are available via Hugging Face (https://huggingface.co/stabilityai). They can be selfhosted, requiring the appropriate computing hardware, including GPUs. DreamStudio—This StabilityAI’s consumer application targets consumers. It is a simple web interface that generates images. The company also has an open source version called StableStudio, driven by the community. More details on DreamStudio can be found at https://dreamstudio.ai. Platform APIs—Stability AI has a platform API (https://platform.stability.ai) that we will use in this book, given that most enterprises would prefer an API that can be managed better at scale. REST API will be used for our example here, as it shows the most flexibility across all platforms. Stable Diffusion also has a gRPC API, which is quite similar. 4.2.1 Dependencies We will build on the packages required earlier in chapter 3 and assume that the following are installed: Python, development IDE, and a virtual environment (such as conda). For Stable Diffusion, we need the following: A Stability AI account and associated API key; this can be acquired via the account page at https://platform.stability.ai/account/keys. Billing details also need to be set up at the same place. We pip install the stability-sdk Python package: pip install stability-sdk. Keep the API key confidential, and follow best practices for managing secrets. We will use environmental variables to store the key securely, which can be configured as follows: –Windows—setx STABILITY API KEY “your-openai-key” –Linux/Mac—export STABILITY API KEY=your-openai-endpoint 110 CHAPTER 4 From pixels to pictures: Generating images –Bash—echo export STABILITY_API_KEY="YOUR_KEY" >> /etc/environment && source /etc/environment We start by getting a list of all the models available using the engines API, including all the available engines (i.e., models). import os import requests import json api_host = "https://api.stability.ai" url = f"{api_host}/v1/engines/list" response = requests.get(url, headers={ "Authorization": f"Bearer {api_key}" }) payload = response.json() # format the payload for printing payload = json.dumps(payload, indent=2) print(payload) The output of this code is presented in the next listing. This shows us the engines we must use and helps in testing end to end to confirm that the API call works and that we can authenticate and get a response. [ { "description": "Real-ESRGAN_x2plus upscaler model", "id": "esrgan-v1-x2plus", "name": "Real-ESRGAN x2", "type": "PICTURE" }, { "description": "Stability-AI Stable Diffusion XL v1.0", "id": "stable-diffusion-xl-1024-v1-0", "name": "Stable Diffusion XL v1.0", "type": "PICTURE" }, { "description": "Stability-AI Stable Diffusion v1.5", "id": "stable-diffusion-v1-5", "name": "Stable Diffusion v1.5", "type": "PICTURE" }, … ] Listing 4.1 Stable Diffusion: Listing the models Listing 4.2 Output: Stable Diffusion model lists REST API call for getting the models HTTP header for authorization Response back from the API Making the JSON more human-readable 1114.2 Image generation with Stable Diffusion 4.2.2 Generating an image We use the Stable Diffusion image generation endpoint (REST API) for our image generation. We will use the latest model, the SDXL model, at the time of this publication. The corresponding engine ID for this model is stable-diffusion-xl-1024-v1-0, as shown in the previous example listing of models. This engine ID is required as part of the REST API path parameter and is available at https://api.stability.ai/v1/generation/ {engine_id}/text-to-image. Listing 4.3 shows an example of using this API to generate an image. Note that we use v1.0 of the API for the examples in this chapter. To use the newer models, we only need to change the REST API path in most cases. For example, to use the newer models that have just been announced, Stable Diffusion 3 and currently in Beta, switch to the following engine ID: https://api.stability.ai/v2beta/stable-image/generate/sd3 . import base64 import os import requests import datetime import re engine_id = "stable-diffusion-xl-1024-v1-0" api_host = "https://api.stability.ai" api_key = os.getenv("STABILITY_API_KEY") prompt = "Laughing panda in the clouds eating bamboo" # Set the folder to save the image; make sure it exists image_dir = os.path.join(os.curdir, 'images') if not os.path.isdir(image_dir): os.mkdir(image_dir) # Function to clean up filenames def valid_filename(s): s = re.sub(r'[^\w_.)( -]', '', s).strip() return re.sub(r'[\s]+', '_', s) response = requests.post( f"{api_host}/v1/generation/{engine_id}/text-to-image", headers={ "Content-Type": "application/json", "Accept": "application/json", "Authorization": f"Bearer {api_key}" }, json={ "text_prompts": [ { "text": f"{prompt}", } ], Listing 4.3 Stable Diffusion: Image generation Choose the model we want to use. Prompts used to generate the image Helper functions to create filenames API call for generating the image The REST API Endpoint includes the engine ID. 118 CHAPTER 4 From pixels to pictures: Generating images As shown in figure 4.15, additional settings for inpainting allow for finer control. Some of these are the same as image generation and are equally important, such as the number of sampling steps and methods. Figure 4.15 Stable Diffusion inpainting options (continued) for that particular task. For example, CLIP can be given the names of visual classes and identify them in images, even if it wasn’t specifically trained on them. CLIP encodes both text and images into a common representation space. It can estimate the most suitable text snippet for an image or vice versa. This gives it much flexibility and the ability to handle different kinds of visual tasks without requiring training data specific to each task. 1194.4 Editing and enhancing images using Stable Diffusion Outpainting is an additional setting that generates and expands the image in our chosen direction. This option is selected via the Script dropdown on the same settings tab (figure 4.16). Figure 4.16 Outpainting settings in Stable Diffusion We go through the iteration of inpainting by removing the areas we want using the mask, regenerating, and then adding the new elements. The final result of these iterations is shown in figure 4.17. NOTE The details on Stable Diffusion web UI, including setup, configuration, and deployments, are outside the scope of this book; however, it is one of the very popular applications that allow one to selfhost across Windows, Linux, and MacOS. You can find more details at their GitHub repository (https:// mng.bz/znx1). 4.4.1 Generating using image-to-image API Image-to-image is a powerful tool for generating or modifying new images that use existing images as a starting point and a text prompt. We can use this API to generate a new image but change the style and mood and add or remove aspects. Figure 4.17 Final edits of inpainting using Stable Diffusion 120 CHAPTER 4 From pixels to pictures: Generating images Let’s use our serene lake example from earlier and then use the image-to-image API to generate a new image. We build on both examples we have seen earlier—we use the serene lake as our input and ask the model to generate “a happy panda eating bamboo in the sky.” import base64 import os import requests import datetime import re engine_id = "stable-diffusion-xl-1024-v1-0" api_host = "https://api.stability.ai" api_key = os.getenv("STABILITY_API_KEY") orginal_image = "images/serene_vacation_lake_house.jpg" #helper functions ... response = requests.post( f"{api_host}/v1/generation/{engine_id}/image-to-image", headers={ "Accept": "application/json", "Authorization": f"Bearer {api_key}" }, files={ "init_image": open(orginal_image, "rb") }, data={ "image_strength": 0.35, "init_image_mode": "IMAGE_STRENGTH", "text_prompts[0][text]": "A happy panda eating bamboo in the sky", "cfg_scale": 7, "samples": 1, "steps": 50, "sampler": "K_DPMPP_2M" } ) data = response.json() for i, image in enumerate(data["artifacts"]): filename = f"{valid_filename(os.path.basename(orginal_image))}_ ➥img2img_{i}_{datetime.datetime.now(). ➥strftime('%Y%m%d_%H%M%S')}.png" image_path = os.path.join(image_dir, filename) with open(image_path, "wb") as f: f.write(base64.b64decode(image["base64"])) We see the generated image as shown on the left in figure 4.18 of the image-to-image API call; we see the panda and the bamboo and how the input image to set the scene Listing 4.4 Image-to-image generation 1214.4 Editing and enhancing images using Stable Diffusion and the type and aesthetic of the generated image are used. However, it doesn’t adhere to the cloud aspect of the prompt. We can tweak the parameters to make it adhere more to the prompt and less to the input image, as shown on the right side of figure 4.18. An example is when we see a panda in the sky, eating bamboo; overall, the image aesthetics follows the input image. Figure 4.18 Stable Diffusion image-to-image generation 4.4.2 Using the masking API Stable Diffusion also has a masking API that allows us to edit portions of an image programmatically. The API is very similar to the creation API, as shown in the example in listing 4.5. It does have a few constraints: the mask image needs to be the same dimension as the original image, and a PNG, less than 4MB in size. The API has the same header parameters outlined earlier in the chapter when we discussed image generation; we will avoid duplicating that. import base64 import os import requests import datetime import re engine_id = "stable-inpainting-512-v2-0" api_host = "https://api.stability.ai" api_key = os.getenv("STABILITY_API_KEY") orginal_image = "images/serene_vacation_lake_house.jpg" mask_image = "images/mask_serene_vacation_lake_house.jpg" prompt = " boat with a person fishing and a dog in the boat" Listing 4.5 Stable Diffusion masking API example Selects the inpainting model we want to use Image we want to edit Masks that we want to apply 122 CHAPTER 4 From pixels to pictures: Generating images # helper functions ... response = requests.post( f"{api_host}/v1/generation/{engine_id}/image-to-image/masking", headers={ "Accept": 'application/json', "Authorization": f"Bearer {api_key}" }, files={ 'init_image': open(orginal_image, 'rb'), 'mask_image': open(mask_image, 'rb'), }, data={ "mask_source": "MASK_IMAGE_BLACK", "text_prompts[0][text]": prompt, "cfg_scale": 7, "clip_guidance_preset": "FAST_BLUE", "samples": 4, "steps": 50, } ) data = response.json() for i, image in enumerate(data["artifacts"]): filename = f"{valid_filename(os.path.basename(orginal_image))}_ ➥masking_{i}_{datetime.datetime.now(). ➥strftime('%Y%m%d_%H%M%S')}.png" image_path = os.path.join(image_dir, filename) with open(image_path, "wb") as f: f.write(base64.b64decode(image["base64"])) Table 4.4 outlines all the API parameters. In terms of options to steer the model, much of it is similar to the previous image creation. Table 4.4 Stable Diffusion masking API parameters Parameter Type Default value Description init_image String Binary (required) The initial image that we want to edit mask_source String Null (required) Mask details that determine the generation areas and associated strengths. It can be one of the following: MASK_IMAGE_WHITE—Use white pixels as the mask; white pixels are modified; black pixels are unchanged. MASK_IMAGE_BLACK—Use black pixels as the mask; black pixels are modified; white pixels are unchanged INIT_IMAGE_ALPHA—Use the alpha channel as the mask. Edit fully transparent pixels, and leave fully opaque pixels unchanged. Masks API call Selects the black pixels of the image to be replaced Prompts for the generation Specifies the number of images to generate Determines the number of steps for each of the images Gets the response from the API Saves the edited image to disk 1234.4 Editing and enhancing images using Stable Diffusion mask_image String Binary (required) Mask image that guides the model on which pixels need to be modified. This parameter is used only if the mask_source is either MASK_IMAGE_BLACK or MASK_IMAGE_WHITE. text_prompts String Null (required) An array of text prompts is used to generate the image. Each element in this array comprises two properties—one of the prompt itself and the other of the associated weight. The weights should be negative for negative prompts. The prompts need to adhere to the following format: text_prompts[index][text|weight], with the index being unique and not having to be sequential. cfg_scale String 7 (optional) Can range between 0 and 35; it defines how strictly the diffusion process follows the prompt. Higher values keep the image closer to the prompt. clip_guidance _preset String None (optional) Different values control how much CLIP guidance is used and influence the quality and relevance of the image being generated. Possible values are NONE, FAST_BLUE, FAST_GREEN, SIMPLE, SLOW, SLOWER, and SLOWEST. sampler String Null (optional) Defines the sampler to use for the diffusion process. If this value is omitted, the API automatically selects an appropriate sampler for you. Possible values are DDIM, DDPM, K_DPMPP_2M, K_DPM_2, K_EULER K_DPMPP_2S_ANCESTRAL, K_HEUN, K_DPM_2_ANCESTRAL, K_LMS, and K_EULER_ANCESTRAL. samples Integer 1 (optional) Defines the number of images to generate. Values need to range between 1 and 10. seed Integer 0 (optional) A random seed is a number that determines how the noise looks. Leave 0 for a random seed value. The possible value ranges between 0 and 4294967295. steps Integer 50 (optional) Defines the number of diffusion steps to run. Possible values range between 10 and 150. style_preset String Null (optional) Used to guide the image model towards a particular preset style. Possible values are 3d-model, analog-film, anime, cinematic, comic-book, digital-art, enhance, fantasy-art, isometric, line-art, low-poly, modelingcompound, neon-punk, origami, photographic, pixel-art, and tiletexture. Note: This list of style presets is subject to change over time. Table 4.4 Stable Diffusion masking API parameters (continued) Parameter Type Default value Description 124 CHAPTER 4 From pixels to pictures: Generating images 4.4.3 Resize using the upscale API The final Stable Diffusion API we want to cover is used to upscale an image, that is, generate a higher-resolution image of a given image. The default is to upscale the input image by a factor of two, with a maximum pixel count of 4,194,304, equivalent to a maximum dimension of 2,048 × 2,048 and 4,096 × 1,024. The API is straightforward, as shown in the next listing. The main thing to be aware of is using the right model via the engine_id parameter. import base64 import os import requests import datetime import re engine_id = "esrgan-v1-x2plus" api_host = "https://api.stability.ai" api_key = os.getenv("STABILITY_API_KEY") orginal_image = "images/serene_vacation_lake_house.jpg" # helper functions ... response = requests.post( f"{api_host}/v1/generation/{engine_id}/image-to-image/upscale", headers={ "Accept": "image/png", "Authorization": f"Bearer {api_key}" }, files={ "image": open(orginal_image, "rb") }, data={ "width": 2048, } ) filename = f"{valid_filename(os.path.basename(orginal_image))}_ ➥upscale_{datetime.datetime.now(). ➥strftime('%Y%m%d_%H%M%S')}.png" image_path = os.path.join(image_dir, filename) with open(image_path, "wb") as f: f.write(response.content) Now that we have examined numerous image-generation options using both GUIs and APIs, let’s examine some of the best practices for enterprises. Listing 4.6 Stable Diffusion resizing API 1254.4 Editing and enhancing images using Stable Diffusion 4.4.4 Image generation tips This section outlines some best practices for image generation. In the context of enterprises, outside of some functions, such as graphic designers and artists, many people with different skills need help. These suggestions will help them get started. We will cover more details later in the book when discussing prompt engineering: Describe in detail—Describe the main subject you want to generate in detail. The visual elements we imagine or want might not match how the model interprets them, so adding details and hints can steer the model more toward what you want. Many also forget to describe the background; it is also important to add those details. Vibes and art style—Specify the style of the vibe or the art that is your intent; for example, we outlined a painting in the prompts earlier. The list is endless and, in some ways, up to your imagination, going from oil painting to steampunk to action photography. Set the emotion, energy, and mood—Add adjectives and verbs that convey the mood, energy, and overall emotion—for example, the generated image aims to be positive and high energy, or positive but low energy, and so forth. Hands and face generations—These are problematic for many models, and while they are getting better, sometimes it is better to add stock or other images to generated images. Structure, size, light, and viewing perspectives—When thinking of the vibe and style of the target image, one also has to think of the size and structure of the artifacts. For example, do we expect something small and intricate or big and freestanding? And from what perspective are the artifacts being looked at—is it a Watermark for AI-generated images Since AI-generated images are getting increasingly better, and we often cannot distinguish between real and AI-generated images, there is a push to watermark AI-generated images. There are two main ways to do this today: visible watermarks, like what Bing and DALLE do, and invisible watermarks, which are not visible to us but are embedded in the image and can be detected using special tools. Google has gone a step further and developed a new type of watermark called SynthID. An invisible watermark is embedded in each image pixel, making it more resistant to image manipulation, such as filters, resizing, and cropping. It does so without degrading the image in any noticeable way and without changing the image size significantly. There are multiple benefits of watermarking AI-generated images. In addition to indicating the origin and possibly ownership of the images, they help discourage unauthorized use and distribution and help prevent the spread of misinformation. Chapter 13 covers GenAI-related risks in more detail, including mitigation strategies and associated tooling. 126 CHAPTER 4 From pixels to pictures: Generating images closeup, a long shot, wide angle, outdoor, or in natural light? Of course, given that we are talking about a prompt, it can combine many of these things. Words, logos, and characters—The image models aren’t large language models and generally struggle with images wherein we expect words to be generated (e.g., a pet salon with its name on the outside). It is best to add these manually when editing the images. Once added, we can use inpainting. Avoid multiple characters together—If you add many characters in the same prompt and generation task, it is common for the model to get confused. It might be better to start with smaller tasks and then use inpainting or manually edit these elements. The next chapter will show other things that can be generated in addition to text and images. We will cover audio, video, and code generators. Summary Vision-based generative AI models allow us to create unique and realistic content, all from a simple prompt. These models can generate new content, edit and enhance existing images, and use simple prompts. Generative AI vision models have multiple use cases in which they can be used for creative content, image editing, synthetic data creation, and generative design. There are four primary generative AI model architectures, each with strengths and challenges. We explained variational autoencoders (VAEs), generative adversarial networks (GANs), vision transformer models (ViT), and diffusion models. Multimodal models are different generative AI models that allow us to handle different types of input data, including text, images, audio, and video, simultaneously. OpenAI’s DALLE, Bing, Adobe, and Stability AI’s Stable Diffusion are some of the more famous and common generative AI image models used by enterprises for image generation and editing. Most things exposed via an API have relevant GUI interfaces too. Many generative AI vision models support inpainting (modifying parts within an image), outpainting (expanding an image beyond its original boundaries), and creating image variations. Diffusion models are more robust in modeling collapse and supporting various outputs. Finally, when it comes to images, we need to think about the scene, main character, structure, and elements such as text and faces, which are better done manually and edited into the image. These aspects have to be added to the prompt for the generation. Later in the book, we will discuss this topic as part of prompt engineering. 127 What else can AI generate? Code that writes itself with little prompting and without much input seems magical, resembling a holy grail, at least to those working in computing. Given the advancements in artificial intelligence (AI) with generative AI, this endeavor seems possible today. We have seen some amazing and interesting things AI can generate—from language to images to holding an ongoing back-and-forth multiturn conversation—and many of them have strong use cases in enterprises. This chapter outlines the remaining things we can generate using AI. We will first talk about code generation, what it means, how one should go about it, and the tools enterprises use. For example, Andrej Karpathy, one of the OpenAI cofounders, who used to lead Tesla’s AI and Vision team, recently said that This chapter covers Using generative AI for code creation and code-related tasks Tools that allow code generation and how to use them Best code generation practices Generating video and related tools Generating audio, music, and related tools 230 CHAPTER 8 Chatting with your data # Get the title ... # get the post description ... # get the publish date ... # get the article body try: article_body = soup.find('div', {'class': 'post-content'}).text except AttributeError: article_body = "" # This should be chunked up article = article_body total_token_count = 0 chunks = [] # split the text into chunks by sentences chunks = split_sentences_by_spacy(article, max_tokens=3000, overlap=10) print(f"Number of chunks: {len(chunks)}") for j, chunk in enumerate(tqdm(chunks)) vector = get_embedding(chunk) # convert to numpy array vector = np.array(vector).astype(np.float32).tobytes() # Create a new hash with the URL and embedding post_hash = { "url": post.link, "title": article_title, "description": article_desc, "publish_date": publish_date, "content": chunk, "embedding": vector } conn.hset(name=f"post:{i}_{j}", mapping=post_hash) p.execute() print("Vector upload complete.") Once we get the blog post’s content, we need to chunk it up, as discussed in the previous chapter. For this example, we use spaCy to chunk the blog post and also have some overlap between different chunks. 8.4.1 Retriever pipeline best practices When implementing a RAG pattern, it’s crucial to have a deep understanding of the source system’s content and structure. The success of a RAG model hinges on its ability to access and interpret the right data, which necessitates a well-architected data 2318.4 Retrieving the data pipeline. This pipeline is not just a conduit for data flow, but a sophisticated framework that ensures data is extracted, transformed, indexed, and stored to align with the model’s requirements and the defined use case. The first step toward implementing GPTs and LLMs in enterprises is a deep understanding of the source system. This involves thoroughly analyzing the data structure, including entity-relationship diagrams, data types, and data distribution. Data profiling tools can be instrumental in understanding the nature of the content. NOTE For RAG to work well, it is important to carefully plan the preprocessing one needs to do in the retriever pipeline and not just use everything without considering whether it is better. If not planned well, this will create problems when using search as part of a RAG implementation. The next phase defines the use case, which entails creating a detailed requirement document outlining the problem, potential solutions, expected results, and success metrics. This document should also detail the users’ informational needs and the scenarios in which the RAG model will be applied. Following this, the focus shifts to data extraction and transformation. This process involves using ETL (extract, transform, load) tools to extract data from the source system and transform it into a format the RAG model can understand. It may involve NLP techniques such as tokenization, stop-word removal, and lemmatization. Once the data has been transformed, it needs to be indexed for efficient retrieval. Azure AI Search, Elasticsearch, Solr, and Lucene are ideal for this purpose, as they provide full-text search capabilities and can handle large datasets effectively. Parallel to data indexing, selecting a suitable data storage solution is important. Depending on the specific needs of the data size, speed, and type, this could be a traditional SQL database, a NoSQL database such as Cosmos DB, or a distributed file system such as Hadoop HDFS. One of the most critical phases is preprocessing planning. This involves careful planning of preprocessing steps, which could involve techniques such as noise removal, normalization, and dimensionality reduction. The goal is to retain information relevant to the use case while reducing the model’s complexity. The next phase is model integration, which involves using APIs or SDKs provided by the AI model vendor to integrate the RAG model into the application. The retriever must be configured with the correct query parameters, and the generator should be set up with the desired output structure. Fine-tuning and monitoring are crucial for enhancing the model’s performance and ensuring the system’s health. This involves using a validation dataset for fine-tuning and application performance management (APM) tools for monitoring. Regarding scalability and reliability, cloud platforms such as AWS, Google Cloud, or Azure should be used to scale the system as needed. Containerization platforms such as Docker and Kubernetes can assist in scaling and managing the application. Redundancy and failover strategies are crucial to ensuring system reliability. 232 CHAPTER 8 Chatting with your data Furthermore, security and compliance cannot be overlooked. Implementing data encryption, user authentication, access control, and regular system audits can ensure data security and compliance with data protection regulations such as GDPR or CCPA. Before deployment, rigorous testing and validation are imperative to ensure that the pipeline and the RAG model meet the expectations outlined by the use case. Once the system is live, comprehensive documentation and technical training should be provided to the team for effective management, maintenance, and troubleshooting. Finally, it’s crucial to ensure the quality control of the retrieval corpus, implement measures for information security and privacy, regularly update the retrieval corpus, and efficiently allocate resources. By following these steps, enterprises can effectively build and maintain AI-powered applications. 8.5 Search using Redis Now that we have the data ingested and the index ready, we can search against it. We create a simple console app that accepts a user’s query, vectorizes it, and searches based on the top three similar posts to return to the user. This is a semantic search. The following listing shows the output generated as an example when we ask about “Longhorn.” $ python .\search.py Connected to Redis Enter your query: Tell me about Longhorn Vectorizing query... Searching for similar posts... Found 3 results: You probably already heard this, but <strong>Chris Sells</strong> ➥has a new column on MSDN called <strong>Longhorn Foghorn</strong> , that describes each of the â <strong>Pillars of Longhorn</strong> â - This is something that IMHO developers would understand and ➥appreciate. In the first article he explains the âPillarsâ and then ➥in the next two goes onto build Solitaire. You can download the sample ➥and play with it too. From OSNews: Microsft has made <em>hard statements about perfomance ➥improvements in Longhorn ... NOTE Windows Longhorn used to be the codename for the operating system that eventually became Windows Vista. Let’s check out the code for implementing the search using Redis. We first take a user query such as “Tell me about Longhorn,” create a vector, and use cosine similarity to obtain a list of comparable results. Listing 8.7 Search results 2338.5 Search using Redis def hybrid_search(query_vector, client, top_k=3, hybrid_fields="*"): base_query = f"{hybrid_fields}=> [KNN {top_k} @embedding $vector AS vector_score]" query = Query(base_query).return_fields( "url", "title", "publish_date", "description", "content", "vector_score").sort_by("vector_score").dialect(2) try: results = client.ft("posts").search( query, query_params={"vector": query_vector}) except Exception as e: print("Error calling Redis search: ", e) return None if results.total == 0: print("No results found for the given query vector.") return None return results # Connect to the Redis server conn = redis.Redis(...) query = input("Enter your query: ") print("Vectorizing query...") query_vector = get_embedding(query) query_vector = np.array(query_vector).astype( np.float32).tobytes() print("Searching for similar posts...") results = hybrid_search(query_vector, conn) if results: print(f"Found {results.total} results:") for i, post in enumerate(results.docs): score = 1 - float(post.vector_score) print(post.content) else: print("No results found") As the name suggests, the hybrid_search() function does the heavy lifting of running the hybrid search query. A hybrid search query combines multiple types of searches into a single query. This can include combining text-based searches with other types, such as numerical, categorical, or even vector-based searches. Note that the exact search type would depend on the information and the requirement. Listing 8.8 Searching using Redis A base query that prefilters fields and is implemented as a KNN search Selects the different fields we are interested in searching Sorts by cosine similarity in descending order Executes the query Captures the query from the user Vectorizes the input Converts the vector to a NumPy array Performs the similarity search 234 CHAPTER 8 Chatting with your data In our example, we combine a K-Nearest Neighbors (KNN) search on an embedding vector with other search fields. The KNN search finds the most related items to a given item, in this case, the most similar posts to a given query vector. The query results are sorted by vector score, which means a high to low ordering based on cosine similarity. In other words, the results with the highest similarity are shown first. We also restrict this to the top three items, as depicted by the top_k parameter. Note that the exact nature of the search and type also depends on the search engine and the data type. For more details on Redis search types and KNN, see the documentation at https://mng.bz/o0Gp. Now that we have seen the search, let’s combine all the dimensions and integrate them into a chat experience using an LLM. 8.6 An end-to-end chat implementation powered by RAG Throughout this and the previous chapter, we have discussed and examined all the pieces to help us understand some of the core concepts; now, we can bring it all together and build an end-to-end chat application. In the application, we can ask questions to get details about our data (i.e., the blog posts). Figure 8.6 shows the application flow. Figure 8.6 End-to-end chat application The question the user asks first gets converted into embeddings and then searched in Redis using a hybrid search index to find similar chunks, which are returned as search results. As we saw earlier, the blog posts have already been injected into the Redis database and indexed. Once we have the results, we formulate the LLM prompt by combining the original questions and the chunks retrieved to answer from. These are passed into the prompt itself before finally calling the LLM to generate a response. Question + Search results Generate answer Question Blog post RSS feed LLM Hybrid search Vector index Create embeddings Redis Formulate prompt Chunks 2358.6 An end-to-end chat implementation powered by RAG On the search front, we deployed Redis running locally and created a vector index. We read all the blog posts going back nearly 20 years. We created the relevant chunks for these posts and their corresponding embeddings and populated our vector database. We also implemented a vector search on those embeddings. The only piece left is to integrate all of this into our application and hook it up with an LLM to complete the last stage of our RAG implementation. Listing 8.9 shows exactly how to do this. Several helper functions, such as get_ search_results(), take the user’s query, call another helper function to search Redis, and return any results found. The actual API call that calls the GPT is in the ask_gpt() function, and it is a ChatCompletion() API, just like we saw earlier. As with previous examples, we leave out the code’s helper functions and other aspects for brevity. The complete code samples are available in the GitHub code repository accompanying the book (https://bit.ly/GenAIBook). def hybrid_search(query_vector, client, top_k=5, hybrid_fields="*"): ... return results def get_search_results(query:str, max_token=4096, ➥debug_message=False) -> str: query_vector = get_embedding(query) query_vector = np.array(query_vector).astype( np.float32).tobytes() print("Searching for similar posts...") results = hybrid_search(query_vector, conn, top_k=5) token_budget = max_token - count_tokens(query) if debug_message: print(f"Token budget: {token_budget}") message = 'Use the blog post below to answer the subsequent ➥question. If the answer cannot be found in the ➥articles, write "Sorry, I could not find an answer in ➥the blog posts."' question = f"\n\nQuestion: {query}" if results: for i, post in enumerate(results.docs): next_post = f'\n\nBlog post:\n"""\n{post.content}\n"""' new_token_usage = count_tokens(message + question + next_post) if new_token_usage < token_budget: if debug_message: print(f"Token usage: {new_token_usage}") message += next_post else: break else: Listing 8.9 End-to-end RAG-powered chat Vectorizes the query Converts the vector to a numpy array Performs the similarity search Manages token budget Loops through the results while still keeping within the token budget 236 CHAPTER 8 Chatting with your data print("No results found") return message + question def ask_gpt(query : str, max_token = 4096, debug_message = False) -> str: message = get_search_results( query, max_token, debug_message=debug_message) messages = [ {"role": "system", "content": "You answer questions in summary from the [CA] blog posts."}, {"role": "user", "content": message},] response = openai.ChatCompletion.create( model="gpt-3.5-turbo-16k", messages=messages, temperature=0.7, max_tokens=2000, top_p=0.95 ) response_message = response["choices"][0]["message"]["content"] return response_message if __name__ == "__main__": # Enter a query while True: query = input("Please enter your query: ") print(ask_gpt(query, max_token=15000, debug_message=False)) print("=="*20) We can see all this coming together when we run it and chat with the blog. It understands the query, creates embeddings, uses the vector database and the associated vector indexes to retrieve the top five matching results, adds that to the prompt, and uses the LLM to generate the response (figure 8.7). In the example we have seen thus far, we are responsible for everything—from setting up the Docker containers to deploying Redis and ingesting the data. This is not enough for enterprises to go into production. More system engineering is required, such as setting up various clusters of machines, scaling them up or down as needed, managing Redis, security requirements, overall operations, and so forth. This takes a significant amount of time, effort, cost, and skills that not every organization might have. Another option is to use Azure OpenAI, which can do much of this out of the box and allows organizations a quicker time to market, potentially at a lower cost. Let’s see how Azure OpenAI can achieve the same result but much faster. Runs a vector search to get embeddings Sets up the chat completion calls Calls the LLM 2378.7 Using Azure OpenAI on your data Figure 8.7 Q&A using blog data with GPT-3.5 Turbo 8.7 Using Azure OpenAI on your data Many enterprises use Azure, and incorporating Azure OpenAI as part of their data strategy represents a pivotal step in employing the power of generative AI for business transformation. Azure OpenAI provides an enterprise-grade platform to integrate advanced AI models such as ChatGPT into your data workflows. “Azure OpenAI on your data” is the service that enables running these powerful chat models on your data and getting out-of-the-box features that enterprises require for production workloads: scalability, security, refreshes, and integration into others. You can connect your data source using Azure OpenAI Studio (figure 8.8) or the REST API. NOTE Azure AI Studio is a platform that combines capabilities across multiple Azure AI services. It is designed for developers to build generative AI applications on an enterprise-grade platform. You can first interact with a project code via the Azure AI SDK and Azure AI CLI and seamlessly explore, build, test, and deploy using cutting-edge AI tools and ML models. At the core of Azure OpenAI’s appeal is its seamless integration with the broader Azure ecosystem. Connecting these powerful AI models to your data repositories unlocks the potential for more sophisticated data analysis, natural language processing, and predictive insights. This integration is particularly beneficial for enterprises with a significant footprint in Azure, enabling them to enhance their existing infrastructure with minimal disruption. 238 CHAPTER 8 Chatting with your data Figure 8.8 Adding your data to Azure OpenAI Azure AI Studio supports multiple options from existing Azure AI Search indexes, Blob storage, Cosmos DB, and so forth. One of these options is a URL, which we will use to ingest blog posts (see figure 8.9). We can also save the RSS feed locally and upload it as a file. One of the advantages of using our own Azure AI Search index is that it does the heavy lifting of keeping the data ingestion up to date from the source systems. This replaces Redis and can be globally distributed to a cloud-scale if required. Figure 8.9 Azure AI Studio: Adding a data source 2398.7 Using Azure OpenAI on your data We can configure and set up most things here, including a storage resource where this data will be saved, an Azure AI Search resource, the index details, embedding details, and so forth (see figure 8.10). With a few clicks, all of this is set up and ready for us to use. Figure 8.10 Configure details for data ingestion On the information security front, this process is streamlined by Azure’s robust security and compliance framework, ensuring that your data remains protected throughout its interaction with AI models. Azure OpenAI supports two key features on your data: role-based and document-level access controls. This feature, working alongside Azure AI Search security filters, can be used to limit access to only those users who should have access based on their permitted groups and LDAP memberships, which is a critical requirement for many enterprises, especially in regulated industries. Finally, Azure’s ability to process and analyze large cloud-scale volumes of unstructured data scalability is another significant advantage. For example, OpenAI’s ChatGPT internally uses Azure AI Search, and that workload is 100+ million users per day. Azure’s cloud infrastructure allows for the easy scaling of AI capabilities as your data needs grow. More details on Azure OpenAI can be found at https://mng.bz/n022. 246 CHAPTER 9 Tailoring models with model adaptation and fine-tuning customer interactions. By using models adapted to their specific needs, businesses can gain insights and increase efficiency, which will provide them with a competitive advantage in their market. Enterprises can enhance efficiency and cost savings by reducing resource requirements and resource needs. Fine-tuning existing models requires significantly less computational power and data compared to training models from the ground up, which results in lower costs and quicker deployment times. For example, training Llama 2’s 70B parameter model took many months and 1,720,320 GPU hours, compared to fine-tuning a GPT-3.5 Turbo model, which takes only a few hours. Model adaptation comes with challenges, and several key areas must be considered. First, task-specific data is crucial. It is essential to have sufficient data to fine-tune an LLM, ensuring that this data is clean, consistent, and representative of the specific task. Depending on the task and LLM characteristics, this data may require preprocessing, augmentation, or labeling. Determining how much data for fine-tuning is enough can be a nuanced process, as it varies based on several factors; at a minimum, it is a few hundred to thousand examples, depending on the model. Determining adequate data for fine-tuning models such as OpenAI’s GPT-3.5 depends on various factors. The complexity and specificity of the task heavily influence data requirements, with more complex tasks requiring more data. However, the quality of data is crucial and often outweighs the quantity. Larger models such as GPT-3.5 can benefit from more data due to their extensive capacity, but they also can learn effectively from smaller, high-quality datasets. Organizations typically start with a baseline dataset and adjust it based on the model’s performance, which is continuously monitored for signs of overfitting or underfitting. Practical constraints such as computational resources and time also play a role in determining the dataset size. The experience and expertise of data scientists often guide the decision. Comparative analysis and continual evaluation are involved in finding the optimal balance of data quantity and quality for the specific task requirements. Another significant challenge is related to computational resources and costs. Fine-tuning LLMs can be resource intensive and costly, often requiring substantial processing power (specifically GPUs) connected with high-speed memory. To manage this, it might be necessary to utilize cloud services, invest in specialized hardware, or employ distributed systems. Additionally, the cost of accessing pretrained LLMs can vary, depending on the provider and licensing agreements, which can add to the overall expense. Performance and generalization are also critical considerations. Evaluating the performance of a fine-tuned LLM is imperative; it involves comparing it to other models or established baselines, which ensures that the fine-tuned LLM does not overfit the training data and can generalize well to new or unseen inputs. We cover evaluations later in this chapter, and more details on benchmarks and associated tools are covered in chapter 12. 2479.2 When to fine-tune an LLM The ethical and social implications of using fine-tuned LLMs must be addressed as well. This includes understanding potential risks and biases, such as concerns related to data privacy, model fairness, and social effects. Adhering to appropriate guidelines, standards, or regulations is necessary to ensure the ethical and responsible use of finetuned LLMs. Finally, finding the right talent is critical. The need for specialized talent and expertise is a significant factor in successfully fine-tuning LLMs, which includes individuals who deeply understand ML, natural language processing (NLP), and the specific architecture of LLMs. These experts must be skilled in various areas, such as data preparation, model architecture design, training strategies, and performance evaluation. The need for skilled personnel adds another layer of challenges to the already complex process of LLM fine-tuning. 9.2 When to fine-tune an LLM Fine-tuning is a technique to improve a model’s performance on a specific task. However, it should be the last option and used only after applying other techniques, such as prompt engineering and RAG. These techniques complement each other and should be stacked for the best output, even when using fine-tuned models. As we saw in earlier chapters, prompt engineering and RAG are not mutually exclusive but are complementary and should be stacked, even when fine-tuning. This stacked approach gives the best outputs, even when using fine-tuned models. Once we decide to fine-tune a model, we prepare the dataset needed for training and start the fine-tuning process, which can take from a few hours to a few days. After training, we evaluate the fine-tuned model against the base model and the specific task’s baseline. Let’s use an example to help us fine-tune and understand various aspects. Say we want to adapt a model to respond with emojis—a bot that can understand what we are asking but respond only using emojis. We will call this EmojiBot. We want to fine-tune GPT-3.5 Turbo and make it an EmojiBot. But to show that these emojis are different and specialized for a task, we don’t want the emojis that we would expect to see, say, in a chat application, on social media, or in our texts. Rather, we want the ones that follow the format used by Microsoft Teams. Figure 9.2 shows the high-level flow for fine-tuning. First, we identify a task that would benefit from fine-tuning (such as EmojiBot). We identify which characteristics fall short of the task and create evaluation criteria. We then compare the default models’ performance against our needs. If they perform well, we establish a baseline and curate the dataset required for fine-tuning. The amount and format of data depend on the model; we’ll cover the details later. We obtain a fine-tuned model after training, which can take hours or days, depending on the task. Next, we must evaluate it against the base model and the baseline for the specific task using qualitative and quantitative measures. 248 CHAPTER 9 Tailoring models with model adaptation and fine-tuning Figure 9.2 Fine-tuning end-to-end flow It is quite common and almost expected that the first fine-tuned model will be worse than the default model. Usually, finding a suitable deployment model takes 10–12 training iterations. Each iteration requires tweaking the training data to address weak areas, which can take hours to days. It’s a timeand effort-consuming process that should be one of the last steps. NOTE Fine-tuning enhances the model’s performance on tasks similar to those outlined in the fine-tuning dataset. This process might manifest as improved accuracy, more relevant responses, or a better understanding of domain-specific language. Improved performance in terms of cheaper or faster models is a side advantage and not something guaranteed. One way to achieve this is to fine-tune a smaller model, such as GPT-3.5 Turbo, on a specific task to improve it instead of using a more expensive and powerful model, such as GPT-4. Now that we have identified a task that makes sense to fine-tune—that is, an EmojiBot where we want to respond in emojis but in a certain pattern—let’s examine the steps needed to fine-tune an LLM such as GPT-3.5 Turbo. 9.2.1 Key stages of fine-tuning an LLM When we want to fine-tune a model for an identified task, as outlined later in figure 9.6, section 9.3.5, there are five key stages: 1Choosing a model and fine-tuning method—To fine-tune a language model, it is necessary to choose a foundation model that suits the task and data. Various Baseline default model Data curation Use cases Task identification Training Production deployment Evaluation No Yes • “EmojiBot” • Evaluation criteria GPT-3.5 Turbo • Training data — Emojis in preferred pattern (sadkoala) • Evaluation data Identify use case fit for fine-tuning Evaluate FT model 2499.3 Fine-tuning OpenAI models models are available, such as GPT, BERT, and RoBERTa. Consider factors such as the model’s suitability for the task, input/output size, dataset size, and technical infrastructure. Fine-tuning methods can vary based on the task and data, such as transfer learning, sequential fine-tuning, or task-specific fine-tuning. 2Data curation—This stage involves preparing a task-specific dataset for finetuning and largely involves preparing and preprocessing the dataset. This process often includes data cleaning, text normalization (e.g., tokenization), and converting the data into a format compatible with the LLM’s input requirements (e.g., data labeling). It is essential to ensure that the data represents the task and domain and covers a range of scenarios the model is expected to encounter in production. 3Fine-tuning—This stage is the actual process of fine-tuning and involves training the pretrained LLM on the task-specific dataset. The training process involves optimizing the model’s weights and parameters to minimize the loss function and improve its performance on the task. The fine-tuning process may involve several rounds of training on the training set, validation of the validation set, and hyperparameter tuning to optimize the model’s performance. 4Evaluating—Once the fine-tuning process is complete, we must evaluate the model’s performance on a test dataset. This helps to ensure that the model is generalizing well to new data and performing well on the specific task. Common metrics used for evaluation include accuracy, precision, recall, F1 score, Bilingual Evaluation Understudy (BLEU), Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and so forth. This topic is covered later in detail in section 9.3.2. 5Deployment (inference)—Once the fine-tuned model is evaluated and we are happy with its performance, it can be deployed to production. The deployment process may involve integrating the model into a larger system, setting up the necessary infrastructure, and monitoring the model’s performance in realworld scenarios. Now that we have a basic concept of model adaptation and when to fine-tune, let’s see how to fine-tune. 9.3 Fine-tuning OpenAI models Here, we’ll use an example to fine-tune OpenAI’s GPT-3.5 Turbo model. Currently, for OpenAI, only GPT-4, GPT-3.5 Turbo, GPT-3 Babbage (Babbage-002), and GPT-3 (Davinci-002) are available for fine-tuning. Several OSS LLMs, such as Meta’s Llama 2 and G42’s Falcon, can be fine-tuned. In our case, the book’s GitHub repository (https://bit.ly/GenAIBook) contains complete code samples and screenshots that we use and show how to fine-tune OpenAI GPT-3.5 Turbo. To make this as real for organizations as possible, we will show the process by using both Azure OpenAI and OpenAI. 250 CHAPTER 9 Tailoring models with model adaptation and fine-tuning We want to fine-tune GPT-3.5 Turbo and make it an EmojiBot, where the model responds in emojis only. However, as we outlined earlier, we want emojis to follow the format used by Microsoft Teams. In Microsoft Teams, the text in parentheses, such as (dog), renders the relevant emojis. We will fine-tune the model to respond to this text, which represents the specific task we want the model to improve. To understand all the different options and the corresponding text in Teams, see https://bit.ly/TeamEmojis. Given that we have a task, let’s start preparing the dataset. 9.3.1 Preparing a dataset for fine-tuning Now that we have reached a point where we have identified a task for which finetuning would make sense, we need to create a dataset of examples required to finetune. We need to create two sets of datasets: one for training and another for validation. A validation dataset is a subset of data used to evaluate the performance of a finetuned model on the target task. It is different from the training dataset, which is used to update the model’s parameters, and the test dataset, which is used to measure the final accuracy of the model. A validation dataset is important for fine-tuning LLMs because it helps us to avoid overfitting, which is when the model learns the specific patterns of the training data and fails to generalize to new data. Using a validation dataset, you can monitor the model’s progress and adjust the learning rate, the number of epochs, or other hyperparameters to optimize the model’s performance. These examples should show different ways to solve the problem and the results of each method. We also need to identify shortcomings using a base model, such as inconsistent performance on edge cases, inability to fit enough shot prompts in the context window to steer the model, high latency, and so forth. It is highly recommended that a validation dataset be used to measure the effectiveness of fine-tuning. The training and validation datasets are in the JSONL format, with each line containing a JSON object with a text key for input text and a target key for desired output text. Fine-tuned models are directly correlated with high-quality training data. Different models require varying amounts of training data. For effective training, we need hundreds to thousands of curated data examples. Although the API requires a minimum of 10 examples, having more is generally better. Ten examples aren’t enough to influence LLMs such as GPT-3.5 Turbo in any significant way. OpenAI recommends having at least 50 good examples to train our model. They also recommend more good examples for better-fine-tuned models than bad ones, as those examples can negatively affect the model. Consequently, it is advisable only to use the best ones from your internal data. The following listing shows an example JSONL file for chat data. 2519.3 Fine-tuning OpenAI models {"messages": [{"role": "system", "content": "You're a chatbot that only responds with emojis!"}, {"role": "user", "content": "I heard the koala habitat was destroyed by the fire."}, {"role": "assistant", "content": "(sadkoala)" }]} {"messages": [{"role": "system", "content": "You're a chatbot that only responds with emojis!"}, {"role": "user", "content": "I've been working nonstop and need a break."}, {"role": "assistant", "content": "(tired)" }]} {"messages": [{"role": "system", "content": "You're a chatbot that only responds with emojis!"}, {"role": "user", "content": "I just finished reading an amazing book!"}, {"role": "assistant", "content": "(like)" }]} As we can see, the model is being shown how to respond using emojis formatted in a certain pattern, such as (sadkoala), (tired), and (like). BASIC CHECKS Before fine-tuning, it’s important to perform basic checks on the training data to avoid wasting time and resources. These checks can include data readability, formatting validation, lightweight analysis for missing pairs, and token length. We validate the data file by loading and reading it using the basic_checks() function. It takes a filename as input and returns the number of messages found. The messages must be in the chat completion format for fine-tuning GPT-3.5 Turbo. # Basic checks to ensure the data file is valid def basic_checks(data_file): try: with open(data_file, 'r', encoding='utf-8') as f: dataset = [json.loads(line) for line in f] print(f"Basic checks for file {data_file}:") print("Count of examples in training dataset:", len(dataset)) print("First example:") for message in dataset[0]["messages"]: print(message) return True except Exception as e: print(f"An error occurred in file {data_file}: {e}") return False FORMAT CHECKS Once we have done the basic checks, the next step is to check the file for the format and ensure it is structured properly before processing it further. This is an important step, mainly because even if the format is incorrect, we won’t get an error when we start the training job, but the resulting model will be very poor, and we will only Listing 9.1 JSONL example Listing 9.2 Dataset validation: Basic checks Opens the file in read-mode Loads each line of the file as a JSON object and stores it in a list Prints the first example from the dataset and helps visually check whether things intuitively look OK Loops through the messages in the first example and prints each one 252 CHAPTER 9 Tailoring models with model adaptation and fine-tuning realize this posttraining when we deploy. To avoid much of this trouble, it is highly recommended that we check for formats. Listing 9.3 shows format_checks(), which checks for chat completion format and pairing, with dataset and filename as its two arguments. It catches most errors but not all. The function iterates over each example in the dataset and checks for data type checks, the presence of message lists, and message keys. It validates that it has the relevant roles and content validation. This function also helps debug data-related problems. def format_checks(dataset, filename): # Initialize a dictionary used to track format errors format_errors = defaultdict(int) # Iterate over each example in the dataset for ex in dataset: # Check if the example is a dictionary, if not # increment the corresponding error count if not isinstance(ex, dict): format_errors["data_type"] += 1 continue # Check if the example has a "messages" key, # if not increment the corresponding error count messages = ex.get("messages", None) if not messages: format_errors["missing_messages_list"] += 1 continue # Iterate over each message for message in messages: # Check if the message has "role" and "content" keys, # if not increment the corresponding error count if "role" not in message or "content" not in message: format_errors["message_missing_key"] += 1 # Check if the message has any unrecognized keys, # if so increment the corresponding error count if any(k not in ("role", "content", "name", ➥"function_call") for k in message): format_errors["message_unrecognized_key"] += 1 # Check if the role of the message is one of the recognized # roles, if not increment the corresponding error count if message.get("role", None) not in ( "system", "user", "assistant", "function", ): format_errors["unrecognized_role"] += 1 Listing 9.3 Dataset validation: Checking for format 2539.3 Fine-tuning OpenAI models # Check if the message has either content or a function call, # and if the content is a string, if not increment the # corresponding error count content = message.get("content", None) function_call = message.get("function_call", None) if (not content and not function_call) or not ➥isinstance(content, str): format_errors["missing_content"] += 1 # Check if there is at least one message with the role "assistant", # if not increment the corresponding error count if not any(message.get("role", None) == "assistant" ➥for message in messages): format_errors["example_missing_assistant_message"] += 1 # If there are any format errors, print them and return False if format_errors: print(f"Formatting errors found in file {filename}:") for k, v in format_errors.items(): print(f"{k}: {v}") return False print(f"No formatting errors found in file {filename}") return True Finally, we should also understand how the dataset performs when it comes to simple data distributions, token counts, and costs. NOTE The token count is important, not just for cost. If it is larger than the maximum number of tokens the model can handle, it will be truncated without warning. Knowing this up front is very helpful. The following listing shows how we can finish doing the checks on the dataset. # Pricing and default n_epochs estimate MAX_TOKENS = 4096 TARGET_EPOCHS = 3 MIN_TARGET_EXAMPLES = 100 MAX_TARGET_EXAMPLES = 25000 MIN_DEFAULT_EPOCHS = 1 MAX_DEFAULT_EPOCHS = 25 def estimate_tokens(dataset, assistant_tokens): # Set the initial number of epochs to the target epochs n_epochs = TARGET_EPOCHS # Get the number of examples in the dataset n_train_examples = len(dataset) # If the examples total is less than the minimum target # adjust the epochs to ensure we have enough examples for Listing 9.4 Dataset validation: Cost estimation and basic analysis 254 CHAPTER 9 Tailoring models with model adaptation and fine-tuning # training if n_train_examples * TARGET_EPOCHS < MIN_TARGET_EXAMPLES: n_epochs = min(MAX_DEFAULT_EPOCHS, MIN_TARGET_EXAMPLES ➥// n_train_examples) # If the number of examples is more than the maximum target # adjust the epochs to ensure we don't exceed the maximum # for training elif n_train_examples * TARGET_EPOCHS > MAX_TARGET_EXAMPLES: n_epochs = max(MIN_DEFAULT_EPOCHS, MAX_TARGET_EXAMPLES ➥// n_train_examples) # Calculate the total number of tokens in the dataset n_billing_tokens_in_dataset = sum( min(MAX_TOKENS, length) for length in assistant_tokens ) # Print the total token count that will be charged during training print( f"Dataset has ~{n_billing_tokens_in_dataset} tokens that ➥will be charged for during training" ) # Print the default number of epochs for training print(f"You will train for {n_epochs} epochs on this dataset") # Print the total number of tokens that will be charged during training print(f"You will be charged for ~{n_epochs * ➥n_billing_tokens_in_dataset} tokens") # If the total token count exceeds the maximum tokens, print a warning if n_billing_tokens_in_dataset > MAX_TOKENS: print( f"WARNING: Your dataset contains examples longer than ➥4K tokens by {n_billing_tokens_in_dataset – ➥MAX_TOKENS} tokens." ) print( "You will be charged for the full length of these ➥examples during training, but only the first ➥4K tokens will be used for training." 9.3.2 LLM evaluation Evaluating LLMs is important for ensuring their quality, reliability, and fairness. However, evaluating LLMs is complex, as it involves multiple dimensions and challenges. Maintaining diverse automatic metrics can help efficiently track model improvements during adaptation cycles, while reducing costly manual reviews. Metrics should be customized to each adapted model’s use cases and business needs. Continuous logging from production systems enables the evaluation of real-world performance over time. Benchmarking against baselines is an essential step in evaluating fine-tuned GPT models. It involves comparing the performance of the fine-tuned model with a preestablished standard or baseline model. This baseline could be the model’s performance before fine-tuning or a different model known for its proficiency in a similar 2559.3 Fine-tuning OpenAI models task. The purpose of this comparison is to quantify the improvements brought by finetuning. For instance, a fine-tuned model might be benchmarked against a standard translation model in a language translation task to assess translation accuracy or fluency improvements. This process helps in understanding the efficacy of fine-tuning and identifying areas where the model has improved or still needs enhancement. EVALUATION CRITERIA When preparing the fine-tuning dataset, we should also define the evaluation criteria. When fine-tuning, the evaluation process begins by establishing clear criteria critical for assessing the performance and efficacy of the model in its intended application. These criteria often include relevance, coherence, accuracy, and language fluency (table 9.1). Evaluating a fine-tuned GPT model using these criteria involves a combination of automated metrics, manual review, and user feedback, ensuring that the model meets the high standards required for its specific application. CHOOSING APPROPRIATE METRICS When fine-tuning models, selecting the right metrics for evaluation is crucial to accurately assessing the model’s performance and improvements [1]. After fine-tuning, these metrics indicate how well the model adapts to specific tasks or domains. They provide insights into various aspects of model performance, such as prediction Table 9.1 Fine-tuning evaluation criteria Evaluation criteria Description Relevance Gauges how well the model’s responses or outputs align with the context and intent of the input. This is especially crucial in applications such as chatbots, where providing contextually appropriate responses is key to user satisfaction. Relevance is often assessed by examining whether the model can stay on topic and provide information or responses directly applicable to the queries or tasks. Coherence Refers to the logical consistency of the model’s outputs. A fine-tuned model should generate contextually relevant, logically sound, and coherent text. This means the responses should follow a logical structure and narrative flow, making sense in the conversation or text context. Coherence is vital for maintaining user engagement and ensuring the model’s outputs are understandable and meaningful. Accuracy This particularly comes into play when the model is used for tasks involving factual information, such as educational tools, informational bots, or any application where providing correct information is critical. Accuracy is measured by how well the model’s responses align with factual correctness and objective truth. Language fluency Pertains to the grammatical and syntactical correctness of the model’s outputs. Even if a model is highly relevant, coherent, and accurate, poor language fluency can significantly detract from the user’s experience. This includes proper grammar, punctuation, and style, ensuring the text generated is correct and reads naturally to the end user. 262 CHAPTER 9 Tailoring models with model adaptation and fine-tuning set, using the same cost function as the training loss. The validation loss is usually measured after each epoch, a complete pass through the training set. Figure 9.3 shows an example of the loss when we fine-tune using Azure OpenAI and the model performance during training. The graph in figure 9.3 showing the training loss for fine-tuning training results illustrates how well the model learns from the training data. We see the loss value for each training step, a batch of training examples. The x-axis is the step number, and the y-axis is the loss value. The graph shows that the loss decreases as the model trains on more data, indicating that it is improving its performance. However, the loss does not reach zero, which means the model still has some errors and cannot perfectly fit the data. This is normal, as overfitting the data can lead to poor generalization of new data. Figure 9.3 Training loss when fine-tuning GPT-3.5 Turbo To interpret the graph and determine whether the model is performing well, ideally for a good fit, we want both training and validation loss to decrease to stability with a minimal gap between the two, which indicates that the model is learning and generalizing well. If the training loss decreases while the validation loss increases, the model may be overfitting the training data and not generalizing well to new data. Finally, if both training and validation loss remain high, the model may be underfitting, which 2639.3 Fine-tuning OpenAI models means it’s not learning the underlying patterns in the data well enough. The scale of the loss and the number of training steps must be considered. The model might need more training if the loss is still high or the validation loss has yet to stabilize. For those with an ML model experience or background, the overall approach for splitting between training and validating datasets and interpreting these metrics is very similar. An interesting behavior is that the data in the loss graph fluctuates, indicating that the loss value can vary depending on the samples in each batch. It is normal for the model to be noisy; however, in fine-tuning, the model learns and improves its performance as long as the loss decreases over time. To find whether the fine-tuning is good, we would typically look for a low and stable validation loss close to the training loss. The thresholds for what would be considered good loss values are subjective and will vary depending on the task’s complexity and the nature of the data. MEAN TOKEN ACCURACY Mean token accuracy measures how well a fine-tuned model correctly predicts each token in the output sequence that the model generates or predicts during training. It is reflected as a percentage, that is, the percentage of tokens the model predicts correctly in a dataset. For example, if the mean token accuracy is 90%, it means that on average, the model correctly predicts 90% of the tokens. This is an average calculated by dividing the number of correctly predicted tokens by the total number of tokens in the output. Similar to the loss for mean token accuracy, we have two metrics: one for the training and the other for validation (assuming one has provided a validation dataset). Figure 9.4 shows the mean token accuracy of a fine-tuning job for training and validation. The training mean token accuracy is the average accuracy of the model’s predictions Figure 9.4 Training mean token accuracy 264 CHAPTER 9 Tailoring models with model adaptation and fine-tuning on the training data. It measures how well the model learns from the training data and adapts to it. A high training mean token accuracy suggests that the model learns effectively from the training data. In contrast, the validation mean token accuracy is the average accuracy of the model’s predictions on the validation data. It measures how well the model generalizes to new data it has not seen before. A high validation mean token accuracy suggests that the model does not overfit the training data and can generalize well to new data. The difference between the two metrics can help identify whether the model is overfitting to the training data. Suppose the training mean token accuracy is much higher than the validation mean token accuracy. In that case, it suggests that the model is overfitting to the training data and not generalizing well to new data. In contrast, if the validation mean token accuracy is much lower than the training mean token accuracy, it suggests that the model is underfitting the training data and not learning effectively. This metric is useful for evaluating the performance of a fine-tuned model on the training data. A good mean token accuracy can be relative and depends on the specific task or application. Generally, a higher value (closer to 1.0) indicates better performance. However, it does not reflect how well the model generalizes to new or unseen data. Note that the interpretation of these metrics can depend on the specific task or application. Therefore, it’s essential to consider other metrics and qualitative evaluations to get a comprehensive view of the model’s performance. The quality of mean token accuracy depends on the task’s complexity and the nature of text. Higher accuracy (closer to 100%) is expected for simpler tasks or texts with predictable patterns. A lower accuracy might still be good for more complex tasks or diverse texts. One way to assess whether the mean token accuracy is good is to compare it with a baseline or with the performance of other models on the same task. If your model’s accuracy is higher than the baseline or similar models, it’s a positive sign. Now that we understand the basic constructs of fine-tuning and using a CLI or code, let’s take a look at how we can achieve this using Azure OpenAI and a GUI. As stated earlier, we will use Azure OpenAI as an example, but the same process applies to OpenAI. 9.3.5 Fine-tuning using Azure OpenAI Instead of using the SDK and the CLI, we also have a visual interface that we can employ to achieve the same outcome. Often, doing this manually would be a better approach than using code. To kick off a fine-tuning job in Azure OpenAI, when logged into the Azure Portal and in the AI Studio, under models, we choose the option to create a custom model (figure 9.5). We go through the wizard and choose to upload the training and validation datasets, as shown in figure 9.6. Note: If these have been uploaded using the SDK, we will find them here, as long as they are in the same tenants and have the same end-point deployment. 2659.3 Fine-tuning OpenAI models Figure 9.5 Azure AI Studio: Creating a custom model Figure 9.6 Choosing a training and validation dataset 266 CHAPTER 9 Tailoring models with model adaptation and fine-tuning Figure 9.7 shows the status and details of each of our training jobs. Now that we have a fine-tuned model, we need to deploy it to a test environment to run an evaluation. 9.4 Deployment of a fine-tuned model The deployment of a fine-tuned model is quite straightforward. The new fine-tuned model shows up as another model available for use in our Azure tenant or OpenAI subscription, as shown in figures 9.8 and 9.9, respectively. Figure 9.7 Training job details Figure 9.8 Deploying fine-tuned model for inference 2679.4 Deployment of a fine-tuned model OpenAI has launched a feature in the playground that lets users see how a fine-tuned model differs from the base model side by side, which can be useful visually but not efficiently. 9.4.1 Inference: Fine-tuned model Returning to our task, we now have a finetuned model for EmojiBot, where the bot responds in emojis using the format that Microsoft Teams uses. Figure 9.10 shows how the out-of-the-box GPT-3.5 Turbo model behaves when asked to respond with emojis; this is expected but will not work with Teams. Figure 9.10 Response with emojis using GPT-3.5 Turbo However, the experience for the same questions using our fine-tuned powered EmojiBot is quite different, as shown in figure 9.11. Here, for the same questions as before, we get the response in the format we’ll be able to use in Teams. Figure 9.9 OpenAI fine-tuned model deployment 268 CHAPTER 9 Tailoring models with model adaptation and fine-tuning Figure 9.11 Fine-tuned EmojiBot inference However, it is easy to get completely incorrect results on the same questions from earlier and with the same parameter settings (figure 9.12). We can see the fine-tuned model answer in emojis—(Pizza) and (Feeling tired)—but the result is not what we expected. Figure 9.12 Fine-tuned EmojiBot with incorrect results 2699.5 Training an LLM To resolve this, we need to tweak the system prompt to steer the model to respond using emojis where possible, which is a great way to close out by reminding that a stacked approach of prompt engineering, RAG, and fine-tuning (where the task at hand warrants) is the right approach in almost all cases. Now that we have seen how to fine-tune a model and the steps one needs to undertake, let us switch and look at some of the underpinnings of the technology that will make this work. Strictly speaking, we do not require this to do a fine-tuning, but it will help us to understand some of the nuances to achieve better outcomes for finetuning. We will start by understanding how we train an LLM and, at a high level, what the steps entail. 9.5 Training an LLM It is helpful to our understanding of model adaptation and the techniques and their associated limitations to examine what it means and what it takes to do full training for an LLM. At a high level, if we were to do full training and build an LLM from scratch, that training would involve four major stages, as shown in figure 9.13. Figure 9.13 Full end-to-end training of an LLM [5] Let’s go through each stage in more detail. 9.5.1 Pretraining Base LLMs are built during this initial stage. We touched on base LLMs in chapter 2. These are the original, pretrained models trained on a massive corpus of text data. They can generate text based on the patterns they learned during training. Some also call them raw language models. Raw internet • Trillions of words • Low quality • Large quantity Stage Pretraining Supervised fine-tuning Reward modeling Reinforcement learning Dataset Algorithm Model Notes Demonstrations • Ideal assistant responses • ~10–100K prompt-response pairs (human written) • High quality; low quantity Comparisons • 100K–1M comparisons (human written) • High quality; low quantity Prompts • ~10K–100K prompts (human written) • High quality; low quantity Language modeling • Next token prediction Language modeling • Next token prediction Binary classification • Predict rewards consistent with preferences Reinforcement learning • Generate token that maximizes the reward Base model SFT model RM model RL model GPUs: Thousands Training: Months Model can be deployed E.g. GPT, PaLM, LLaMA GPUs: 1–100 Training: Days Model can be deployed E.g. Vicuna-13B GPUs: 1–100 Training: Days GPUs: 1–100 Training: Days Model can be deployed E.g. ChatGPT, Claude Initialized Initialized Initialized Use RM 270 CHAPTER 9 Tailoring models with model adaptation and fine-tuning NOTE While powerful, these base models are less suitable for generalpurpose applications because they may need to align their responses with the specific intentions or instructions of the user. They are more like raw engines for text generation, lacking the refined capability to understand and adhere to the nuances of user prompts. Base models do not answer questions and often respond with more questions. In contrast, instructors are tailored to be more interactive and user-friendly, which makes them more suitable for a wide range of applications, from customer service chatbots to educational tools, where understanding and following instructions accurately is crucial. 9.5.2 Supervised fine-tuning Supervised fine tuning (SFT) is the next stage. In this stage, the base model undergoes refining of the base model with high-quality, domain-specific data. These datasets consist of prompt–response pairs, manually created (often by human contractors), which are fewer in number than in the previous stage but of much higher quality. The contractors follow detailed documentation to create these prompt–response pairs, ensuring relevance and quality. Similar to the last pretraining stage, the SFT model is trained to predict the next token in these pairs, but these are less accurate and contextually aware when generating the response. SFT is a technique for optimizing LLMs on labeled data for a specific downstream task, such as sentiment analysis, text summarization, or machine translation. Later in the chapter, we will cover additional details of SFT methods and approaches. 9.5.3 Reward modeling The third phase is reward modeling, the first part of the Reinforcement Learning from Human Feedback (RLHF) process. The main goal at this stage is to develop a model that can evaluate and rank responses based on their quality and relevance. To do this, the SFT model (from the previous stage) generates multiple responses to a prompt, which human contractors then rank based on various criteria such as domain expertise, fact-checking, and code execution. These rankings train a reward model, which learns to score responses like human contractors. 9.5.4 Reinforcement learning This is the second part of the RLHF process, and it aims to enhance the language model’s ability to generate high-quality responses through iterative feedback. In this final stage, the reward model scores responses generated by the SFT model for many prompts. These scores are used to further train the SFT model, ultimately leading to the creation of the RLHF model. The RLHF aligns the LLMs with human preferences or expectations for a given task or domain, such as chat, code, or creative writing. More details on RLHF methods will be covered later in this chapter. 9.5.5 Direct policy optimization Direct policy optimization (DPO) [6] is another technique, which is a new type of reward model parameterization in RLHF that used for fine-tuning LLMs to align with 2719.6 Model adaptation techniques our preferences. It exploits a relationship between reward functions and optimal policies. It allows us to skip the reward modeling step outlined earlier, as long as the human feedback can be expressed in binary terms—that is, a choice between two options. DPO can solve the reward maximization problem with constraints in a single policy training phase, essentially treating it as a classification problem. PPO (see section 9.7) requires a reward model and a complex RL-based optimization process; DPO, however, bypasses the reward modeling step and directly optimizes the language model on preference data, which can be simpler and more efficient. As DPO eliminates the need to train a reward model instead of training a reward model and optimizing a policy based on that model, we can directly optimize the policy. This characteristic makes this approach quicker, and fewer resources are used than in RLHF with PPO. 9.6 Model adaptation techniques There are several techniques available for model adaptation, with each technique providing its unique approach and being suitable for different scenarios depending on the specific requirements (i.e., the model size, available computational resources, and the desired level of adaptation). One of the main techniques widely used for adapting LLMs is low adaptation ranking (LoRA), which will be covered in more detail in the next section. In LoRA, instead of updating all the weights in the model, only a small subset of parameters, introduced as low-rank matrixes, are modified. This approach allows efficient training and adaptation, while preserving most of the pretrained model’s structure and knowledge. Parameter efficient fine-tuning (PEFT) is a concept in ML that refers to methods of adapting and fine-tuning large pretrained models, such as GPT-3.5, to minimize the number of parameters that need to be updated. This approach is particularly valuable when dealing with large models, as it reduces computational requirements and can mitigate problems such as overfitting. PEFT techniques are designed to make fine-tuning more accessible and efficient, especially for users with limited computational resources—LoRA is an example of the PEFT method. For more details on different types of PEFT techniques and details, see the paper “Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning” by Vladislav Lialin [7]. Catastrophic forgetting is a phenomenon where a model loses its ability to perform well on previous tasks after being fine-tuned on new tasks [8]. This can happen when the model overwrites its original parameters with task-specific ones, thus forgetting the general knowledge it learned from pretraining. When implementing PEFT to prevent catastrophic forgetting, we fine-tune only a small subset of parameters, while keeping most pretrained parameters fixed. This way, the model can retain its generalization ability and adapt to new tasks without losing its previous performance. Supervised fine-tuning (SFT) is another type of adaptation technique; it is a specific type of fine-tuning where the model is further trained on a labeled dataset. It’s supervised because the training process uses a dataset that pairs the input data with the correct output (labels). SFT is particularly common in tasks such as classification, 278 CHAPTER 9 Tailoring models with model adaptation and fine-tuning Proximal policy optimization (PPO)—PPO is an RL algorithm that iteratively improves the primary model’s policy (decision-making process). The algorithm updates the model’s policy to maximize the rewards predicted by the reward model. PPO is chosen for its stability and efficiency in handling large and complex models. Human feedback loop—This loop involves continuous input from human evaluators who assess the quality of the model’s outputs. The feedback is used to train the reward model further, creating a dynamic learning environment where the model adapts to evolving human preferences and standards. The loop ensures that the model remains aligned with human expectations and can adapt to changes. NOTE PPO-ptx [13] is an adaptation of the PPO algorithm tailored for finetuning RLHF. It integrates a reference to the original LLM to maintain performance, while aligning the model’s outputs with human preferences. This approach helps mitigate the alignment tax, ensuring the LLM remains effective and diverse in its outputs after training. Essentially, PPO-ptx balances the model’s pretraining knowledge with the new feedback to create a highperforming LLM aligned with human values. RLHF might seem like the silver bullet in many ways, but enterprises must be aware of some challenges. Let’s explore these. 9.7.1 Challenges with RLHF RLHF is a powerful technique for teaching models complex tasks, but it has many practical challenges and limitations for enterprises. An RLHF system needs a lot of human preference data, which is hard to get because it involves other people who are not part of the training process. How well RLHF works depends on how good the human annotations are, which humans can write, such as when they adjust the initial LLM in InstructGPT or provide ratings of how much they like different outputs from the model. Some of these challenges are Technical complexity—Implementing RLHF requires advanced skills and knowledge in ML, RL, and NLP. It also involves complex setup and maintenance processes, such as configuring the model architecture, reward systems, and feedback mechanisms. Computationally intensive—RLHF models need a lot of computational resources, such as GPUs and servers, which can be expensive. They also depend on the quality and quantity of human feedback, which can be hard to obtain and process. From a practical viewpoint, a lot of the human feedback is from contract workers (or gig workers) on crowdsourcing platforms where getting the right qualified people in certain domains might be challenging. Moreover, ensuring a diverse and unbiased dataset for training can be challenging and computationally heavy. 2799.7 RLHF overview Not scalable—RLHF models are difficult to scale for large-scale applications, requiring continuous human feedback and increasing computational resources. They are also hard to adapt to different domains or changing data environments, resulting in limited adaptability and customization. Quality—RLHF models are prone to bias, as they reflect human feedback providers’ subjective opinions and potential prejudices. Ensuring ethical use and unbiased outputs is a major concern. Maintaining a consistent quality of human feedback can be difficult, as human judgment can vary and affect the model’s reliability and performance. When trying to build a helpful model that avoids harm, there is an inherent tension between those two dimensions. Providing too many polite responses, such as “Sorry, I am an AI model, and I cannot help you with that,” or something similar, limits the model’s usefulness. Organizations must balance and mitigate this using additional guidance, training, and other ML techniques to create synthetic data where possible. Cost—RLHF models are costly to implement and operate. Costs include infrastructure, computational resources, data acquisition, and hiring skilled professionals. There are also ongoing operational costs related to data management, model updates, and continuous feedback integration. These costs can be substantial, especially for large-scale implementations. Data—It is hard to produce good human text that answers specific prompts because it usually means paying part-time workers (instead of product users or crowdsourcing). Luckily, the amount of data needed to train the reward model for most uses of RLHF (~50k preference labels) is not that costly. However, it is still more than what academic labs can usually afford. There is only one big dataset for RLHF on a general language model (from Anthropic) and a few smaller datasets for specific tasks (such as summarization data from OpenAI). Another problem with data for RLHF is that human annotators can disagree a lot, which makes the training data very noisy without a true answer. RLHF offers advanced capabilities in teaching models to perform complex tasks; however, its adoption in enterprise settings is hindered by technical complexity, resource demands, scalability challenges, ethical considerations, and high costs. On the one hand, these barriers make it difficult for many organizations to implement and sustain RLHF systems in their operations practically. On the other hand, those who can implement this, especially some of the technical companies such as OpenAI and Anthropic, can benefit from it. Let’s see how we can scale an RLHF implementation. 9.7.2 Scaling an RLHF implementation Scaling an RLHF implementation for LLMs involves a multifaceted approach that balances efficiency, diversity, and quality control. First, automating data collection and implementing efficient feedback mechanisms are crucial for handling large volumes of data and feedback. Automated systems can gather data from various sources or through interfaces designed for efficient human interaction. 280 CHAPTER 9 Tailoring models with model adaptation and fine-tuning Using a large, diverse pool of human evaluators is essential for capturing a wide range of perspectives, helping the model to be more robust and less biased. To ensure the feedback is informative, intelligent sampling strategies, such as active learning, can be used to identify and prioritize the most valuable instances for evaluation. Parallelization and distribution of tasks among multiple evaluators can significantly speed up the feedback process. The system can handle large-scale data processing and model training with scalable infrastructure. Implement quality control measures, such as cross-validation among evaluators and algorithms, to detect biases and maintain the quality and consistency of feedback. Regular monitoring and evaluation of the model’s performance can help you understand the effects of RLHF and guide continuous improvement. Finally, ethical considerations and bias mitigation are crucial. Ensuring that feedback does not reinforce harmful stereotypes and actively addressing potential biases is vital for developing fair and responsible models. Overall, scaling RLHF for LLMs requires a comprehensive approach that integrates technical, logistical, and ethical strategies, aiming for a system that effectively incorporates human feedback into the model’s learning process. Summary Model adaptation should be anchored in a set of use cases, and it should be the last resort for enterprises trying to improve the model on those tasks. Prompt engineering and RAG must work in conjunction with fine-tuning in a stacked manner. When done correctly, fine-tuning has a high upside from enhanced efficiency and possible cost savings. Fine-tuning has a high cost, and you should be aware of challenges such as the need for task-specific data, computational resources, performance evaluation, and ethical considerations. Fine-tuning should be done in conjunction with evaluations and will often require multiple iterations to obtain a model ready for production deployment. The choice of metrics for evaluating fine-tuned models largely depends on the model’s specific application and objectives. The main model adaptation techniques that are more cost-efficient are supervised fine-tuning (SFT), parameter efficient fine-tuning (PEFT), and low-rank adaptation (LoRA). Part 3 Deployment and ethical considerations This final section focuses on the practical aspects of deploying generative AI applications and the ethical considerations involved. It provides a comprehensive guide to application architecture, scaling up for production, and the operational best practices for deployment. The closing chapters emphasize the importance of ethical principles, discussing potential risks, responsible AI lifecycle, and tools for ensuring ethical AI practices. Chapter 10 discusses the architectural considerations necessary for building generative AI applications. It covers the orchestration and grounding layers and how to filter models and responses to effectively ensure optimal application performance. Chapter 11 focuses on the challenges of scaling generative AI applications and provides best practices for production deployment. It addresses critical aspects such as metrics, latency, scalability, and security considerations to ensure smooth and efficient operation. Chapter 12 explains how to evaluate and benchmark large language models, discussing various metrics and benchmarks. It covers task-specific benchmarks and the importance of human evaluation in assessing model performance. Chapter 13, the final chapter, highlights generative AI’s ethical challenges and risks. It outlines the principles and practices for responsible AI use, including content safety, data privacy, security considerations, and the ethical lifecycle of AI implementation. 282 CHAPTER 283 Application architecture for generative AI apps The enterprise architecture landscape continues to change, moving inexorably toward more self-directed systems—intelligent, self-managing applications that are capable of learning from interactions and adapting in real time. Furthermore, increasing digitization fuels the AI digital transformation. This ongoing progression This chapter covers An overview of GenAI application architecture and the emerging GenAI app stack The different layers that make up the GenAI app stack GenAI architecture principles The benefits of orchestration frameworks and some of the popular ones Model ensemble architectures How to create a strategic framework for a crossfunctional AI Center of Excellence 284 CHAPTER 10 Application architecture for generative AI apps underscores a transformative era in enterprise technology, poised to redefine the very nature of software development and deployment. Naturally, this is more of an ideal. However, most enterprises are still very inexperienced with AI-infused applications in general, and generative AI is still very much in its early stages. This chapter will explore how enterprise application architecture standards and best practices must adapt to the emerging generative AI technologies and use cases. The chapter introduces the concept of a GenAI app stack as a conceptual reference architecture for building generative AI applications, and it outlines its main components and how generative AI fits together in the broader enterprise architecture. The GenAI app stack is an evolution of cloud application architecture, with a shift toward data-centric and AI-driven architectures. This chapter starts by outlining what the new GenAI app stack entails, covering details of each section and, finally, bringing all the concepts together into working examples that make it real and usable. As you learn about this stack, we’ll consolidate the different aspects of the architecture described in previous chapters. One thing to note is that despite representing a big change, generative AI does not require a completely new architecture but builds on the existing cloud-based distributed architecture. This characteristic allows us to build on existing best practices and architecture principles to incorporate new GenAI-related paradigms. Let’s start by identifying the updates to enterprise application architecture. 10.1 Generative AI: Application architecture Over the last few years, enterprise application architecture has witnessed a significant evolution, going through several transformative stages to meet the escalating demands for business agility, scalability, and intelligence. Initially, enterprises operated on monolithic systems, that is, robust but inflexible structures with tightly interwoven components, which made changes cumbersome and wide-reaching. These systems set the stage for enterprise computing but were not suitable for the rapid evolution of business needs. The proliferation of cloud computing and cloud-native architectures saw the rise of containerization and orchestration tools, which simplified the deployment and management of applications across diverse environments. Simultaneously, the deluge of data led to data-centric architectures that prioritize data processing and analytics as key drivers for business operations. The evolution of enterprise application architecture for generative AI can be seen as a shift from traditional software development to data-driven software synthesis. In the traditional paradigm, software engineers write code to implement specific functionalities and logic, using frameworks and libraries that abstract away low-level details. In the generative AI paradigm, software developers provide data and highlevel specifications and use large language models (LLMs) to generate code that meets the desired requirements and constraints. The following two key concepts enabled this paradigm shift: Software 2.0 and building on copilots. 28510.1 Generative AI: Application architecture 10.1.1 Software 2.0 Software 2.0 is a term coined by Andrej Karpathy [1] to describe the trend of replacing handcrafted code with learned neural networks. Software 2.0 uses advances in AI, such as natural language processing (NLP), computer vision, and reinforcement learning, to create software components that can learn from data, adapt to new situations, and interact with humans naturally. Recently, we have transitioned from writing code and managing explicit instructions for a desired goal to a more abstract approach. Developers train models on large datasets instead of writing explicit instructions or rules in a programming language. Software 2.0 also reduces the need for manual debugging, testing, and maintenance, as the neural networks can self-correct and improve over time (see figure 10.1). Figure 10.1 Software 1.0 versus Software 2.0 This allows the models to learn the rules or patterns themselves. Algorithms and models are crafted to learn from data, make decisions, and improve over time, effectively writing the software. This paradigm shift has transformed the role of AI from a supportive tool to a fundamental component of system architecture. 10.1.2 The era of copilots Another key concept that facilitated the evolution of enterprise application architecture for generative AI is copilots—a concept originally proposed by Microsoft. Copilots are meant to augment humans and human capabilities and creativity. Using an 30 Input data func foo(x): int { return x+1 } Code Computation 6 Output Computation Weights Output Labelled training data Model architecture Software 1.0 Software 2.0 286 CHAPTER 10 Application architecture for generative AI apps airplane analogy, if we are humans, we are the pilots; instead of AI being on autopilot where we have no control or say in how it functions, this new AI plays the role of copilots that help us take on cognitive load and some of the drudgery of work. Still, we remain in charge as the pilot. The Copilot stack is a framework for building AI applications and copilots that use LLMs to understand and generate natural language and code. Copilots are intelligent assistants that can help users with complex cognitive tasks such as writing, coding, searching, or reasoning. Microsoft has developed a range of copilots for different domains and platforms, such as GitHub Copilot, Bing Chat, Dynamics 365 Copilot, and Windows Copilot. You can also build your custom Copilot using the Copilot stack and tools, such as Azure OpenAI, Copilot Studio, and the Teams AI Library. Copilots can also be integrated into existing tools and platforms, such as GitHub, Visual Studio Code, and Jupyter Notebook, to enhance the productivity and creativity of software developers. Copilots are based on the concept of Software 2.0, where they use LLMs to generate code from natural language descriptions instead of relying on manually written code. However, they should be seen as the GenAI application stack, similar to the LAMP stack for web development. LAMP is an acronym for the stack components: Linux (operating system); Apache (webserver); MySQL (database); and PHP, Perl, or Python (programming language). Copilots are a useful model for enterprises to follow when designing their generative AI apps enterprise architecture because they offer several advantages (e.g., quicker and simpler development, more creativity and testing, and improved cooperation and learning, enabling enterprises to try out new concepts and opportunities or to create original solutions for difficult problems). Let’s expand on what the Copilot stack is to make it more relevant and real in concrete terms. 10.2 Generative AI: Application stack Copilots’ architecture comprises several layers and components that work together to provide a seamless and powerful user experience, as outlined in figure 10.2. We will start from the bottom up, examine each layer and component in detail, and find out how they interact. The AI infrastructure layer is the foundational layer that powers everything and hosts the core AI models and computational resources. It encompasses the hardware, software, and services that enable the development and deployment of AI applications and are often optimized for AI workloads. This also includes the massively scalable distributed high-performance computing (HPC), required for training the base foundational models. The foundational model layer includes the range of supported models, from hosted foundation models to the model you train and want to deploy. The hosted foundational models are large pretrained models, such as LLMs and others (vision and speech models), and the newer small language models (SLMs) that can be used for inference; these models can be closed or open. Some of the models can be further 28710.2 Generative AI: Application stack adjusted for specific tasks or domains. These models are hosted and managed within the AI infrastructure layer to ensure high performance and availability. Users can select from various hosted foundation models based on their needs and preferences. The orchestration layer manages the interactions between the various components of the architecture, ensuring seamless operation and coordination. It is responsible for key functions such as task allocation, resource management, and workflow optimization: The response filtering component uses the prompt engineering set of components; here, the prompts and responses are analyzed, filtered, and optimized to generate safe outputs. The system prompt can also provide additional information or constraints for the AI model to follow. The user can express a system prompt via a simple syntax, or the system can automatically generate it. Grounding is the implementation of retrieval-augmented generation (RAG), and it refers to the process of contextualizing the responses generated by the AI model. Grounding ensures the outputs are syntactically correct, semantically meaningful, and relevant to the given context or domain. We use plugins to get data ingested from different enterprise systems. The plugin execution layer runs plugins that add more features to the basic AI model. Plugins are separate and reusable parts that can do different things, AI infrastructure Foundational models Hosted foundational models Hosted fine-tuned foundational models BYO models Grounding Plugin execution VectorDB, APIs, etc.) ( Response filtering Meta prompt Orchestration Copilot frontend + UX AI safety Figure 10.2 GenAI application stack 294 CHAPTER 10 Application architecture for generative AI apps getting bogged down in the technical details of LLM interaction. Table 10.1 outlines the key responsibilities. These different components work together to create a strong orchestration system that serves as the foundation for the successful deployment and operation of generative AI technology in the enterprise sector. Such orchestration is necessary for the intricacy and constant changes of AI-powered applications to avoid inefficiencies, mistakes, and system breakdowns. 10.3.1 Benefits of an orchestration framework Orchestrators are essential for managing the complex systems powering generative AI apps. These systems involve diverse processes that need careful coordination through orchestration tools. Orchestrators simplify workflows and ensure tasks are done in Table 10.1 Orchestrator key responsibilities Area Descriptions Workflow management Orchestrator ensures that the sequence of processes—from data ingestion and processing to AI model inference and response delivery—is executed in an orderly and efficient manner. This includes state management to coordinate dependencies between tasks, error handling, retry mechanisms, and the dynamic allocation of resources based on the task load. Service orchestration Microservices architecture is typically employed, where each service is responsible for a discrete function in the generative AI process. Service orchestration is about managing these services to scale, communicate, and function seamlessly. In addition, containerization platforms such as Docker and orchestration systems such as Kubernetes deploy, manage, and scale the microservices across various environments. Data flow coordination Ensure that data flows correctly through the system, from the initial data sources to the model and back to the end user or application. This includes preprocessing inputs, queue management for incoming requests, and routing outputs to the correct destinations. Load balancing and auto-scaling Load balancers distribute incoming AI inference requests across multiple instances to prevent any single instance from becoming a bottleneck. Autoscaling adjusts the number of active instances based on the current load, ensuring cost-effective resource use. This also has API management components to manage rate limits and implement back-off strategies for production workloads. Model versioning and rollback Orchestration includes maintaining different versions of AI models and managing their deployment. It allows for quick rollback to previous versions if a new model exhibits unexpected behavior or poor performance. Managing model context windows Orchestrator enhances interactions by efficiently managing context windows and token counts. It tracks and dynamically adjusts conversation history within the model’s token limits and maintains coherence in responses, especially in long or complex exchanges. Best practices include efficient context management, handling edge cases, continuous performance monitoring, and incorporating user feedback for ongoing improvements. 29510.3 Orchestration layer order, with dependencies and error-handling rules taken care of. This results in a reliable and regular operational flow, where steps for preprocessing, computation, and postprocessing are smoothly connected, ensuring data quality and consistent output generation. Scalability is another area where orchestration is vital. As demand fluctuates, a system that dynamically adjusts resource allocation, especially for production workloads, becomes crucial. An orchestrator can provide this agility using different techniques, such as load balancers to distribute workloads evenly and auto-scaling features to modulate computing power in real-time. This elasticity meets the load requirements and optimizes resource usage, balancing performance and cost efficiency. The orchestrators would need to manage this across different models, as well as the computational and cost profiles of those models. Orchestrators offer a centralized management and monitoring ability. They constitute frameworks that offer dashboards and tools for monitoring LLM usage, identifying bottlenecks, and troubleshooting problems. This enhances system reliability by monitoring service health, responding to failures, and ensuring minimal downtime. Orchestrators can employ automated recovery processes, such as instance restarts or replacements, allowing for service continuity. The default deployment model is a pay-as-you-go method for most cloud-based LLM providers. This model is shared with other customers, and incoming requests are queued and processed on a first-come, first-served basis. However, for production workloads that require a better user experience, Azure OpenAI service offers a provisioned throughput units (PTU) feature. This feature allows customers to reserve and deploy units of model processing capacity for prompt processing and generating completions. Each unit’s minimum PTU deployment, increments, and processing capacity vary depending on the model type and version. An orchestrator will manage the different deployment endpoints between regular pay-as-you-go and PTUs to ensure optimum performance and cost-effectiveness. Orchestrators play a significant role in increasing productivity and streamlining operations, which are achieved in two ways. First, it reduces the need to write repetitive code for common tasks such as prompt construction and output processing, thus increasing developers’ productivity. Second, it automates the deployment and management of services, thus minimizing the possibility of human error. This automated process reduces manual overhead and ensures effective compute resource utilization, streamlining production operations. We will delve deeper into managing operations later in the chapter. Compliance and governance are essential requirements for any enterprise. An orchestrator can assist in enforcing compliance by determining how data is processed, stored, and used in the workflow, which ensures that the data complies with the enterprise’s data governance policies and privacy regulations. Maintaining trust and legal compliance in enterprise operations is crucial and can be achieved through adherence to data governance policies and privacy regulations. 296 CHAPTER 10 Application architecture for generative AI apps 10.3.2 Orchestration frameworks Many people are familiar with orchestrators and orchestration frameworks. While frameworks such as Kubernetes, Apache Airflow, and MLflow are effective general orchestration tools for software engineering and can support ML operations, they are not designed exclusively for generative AI applications. Orchestrating workflows for generative AI requires a more intimate understanding of the nuances of these complex technologies. The choice of an orchestration framework for generative AI applications depends on the existing technology stack, the complexity of the workflows, and specific requirements. Table 10.2 outlines orchestration frameworks tailored to the specific needs of generative AI applications. These frameworks can handle traditional computational workflows; manage interactions’ state, context, and coherence; and are designed to suit the unique requirements of generative AI. Table 10.2 Orchestration frameworks Name Notes Semantic Kernel Semantic Kernel is an OSS framework from Microsoft that aims to create a unified framework for semantic search and generative AI. It uses pretrained LLMs and graph-based knowledge representations to enable rich and diverse natural language interaction. LangChain LangChain is a library that chains language models with external knowledge and capabilities. It facilitates the orchestration of LLMs such as GPT-4 with databases, APIs, and other systems to create more comprehensive AI applications. PromptLayer PromptLayer is a platform that simplifies the creation, management, and deployment of prompts for LLMs. Users can visually edit and test prompts, compare models, log requests, and monitor performance. More details can be found at https://promptlayer.com/. Rasa Rasa is an enterprise conversational AI platform that lets you create chatand voice-based AI assistants to manage various conversations for different purposes. In addition to conversation AI, it also offers a generative AI-native method for building assistants, with enterprise features such as analytics, security, observability, testing, knowledge integration, voice connectors, and so forth. More information is available at https://rasa.com/. YouChat API The YOU API is a suite of tools that helps enterprises ground the output of LLMs in the most recent, accurate, and relevant information available. You can use the YOU API to access web search results, news articles, and RAG for LLMs. More details can be found at https://api.you.com/. Ragna Ragna is an open source RAG-based AI orchestration framework that allows you to experiment with different aspects of a RAG model—LLMs, vector databases, tokenization strategies, and embedding models. It also allows you to create custom RAG-based web apps and extensions from different data sources. More details can be found at https://ragna.chat/. LlamaIndex LlamaIndex is a cloud-based orchestration framework that enables you to connect your data to LLMs and generate natural language responses. It can access various LLMs. Hugging Face Hugging Face provides a collection of pretrained models for various NLP tasks. It can be used with orchestration tools to manage the lifecycle of generative AI applications. More details can be found at https://huggingface.co/. 29710.3 Orchestration layer 10.3.3 Managing operations An orchestrator plays a crucial role in enhancing the performance and seamless integration of generative AI models, such as LLMs, within intricate systems and workflows. Its core functionality optimizes operational efficiency and fosters a better user experience through sophisticated control mechanisms. The orchestrator is crucial in managing the LLM’s integration into complex workflows, such as content creation pipelines. It plans and schedules the LLM’s activation to ensure smooth data collection, preprocessing, and text generation, thus simplifying the entire process from start to finish. This coordination improves the workflow and ensures that the API calls for the generated content are timely and relevant. The orchestrator’s main role is to balance the load and resources for the LLM’s services. It effectively manages requests to avoid overloading or wasting resources. Furthermore, it can change computational resources by constantly tracking workload and performance metrics. This flexibility ensures the system stays responsive and resources are used efficiently, even when demanding changes. The orchestrator also supervises API interactions, enforcing rate limits and controlling secure access, while managing any errors or disruptions that may occur. Simultaneously, it handles the essential tasks of data preprocessing and postprocessing. This means cleaning, formatting, and transforming data to ensure it is in the right state for processing by the LLM and then improving the output to meet set quality standards and format requirements. For workflows requiring sequential processing, the orchestrator ensures that outputs from one phase are accurately fed into the next, maintaining the process integrity. This is complemented by its role in enforcing security and compliance measures, where it filters sensitive information and ensures adherence to legal and ethical standards, in addition to conducting audits for accountability and quality assurance. For applications such as chatbots or digital assistants, the orchestrator manages user interactions by handling session states and queries, directing them to the LLM or other services as needed, which results in a more engaging and responsive user experience. Moreover, the orchestrator continuously monitors the LLM performance, analyzing response time, accuracy, and throughput to guide optimization efforts. It also manages updates to the LLM, ensuring that transitions to newer versions or configurations are smooth and minimally disruptive to users. As we can see, an orchestrator can significantly enhance the efficiency, reliability, and scalability of an LLM when integrated into complex systems, providing a layer of management that coordinates between the LLM and other system components. Building your own orchestrator framework Creating your own generative AI orchestrator for an enterprise can be difficult. However, it allows you to customize the framework according to your requirements and increases your understanding of the technology. This process demands extensive 298 CHAPTER 10 Application architecture for generative AI apps Some new frameworks used widely nowadays are Semantic Kernel, LangChain, and LlamaIndex. These frameworks enable the use of GenAI models, although they address different aspects. We will explore these in more depth. SEMANTIC KERNEL Semantic Kernel (SK) from Microsoft is an SDK that integrates LLMs with languages such as C#, Python, and Java. It simplifies the sometimes-complex process of interfacing LLMs with traditional C#, Python, or Java code. With SK, developers can define semantic functions that encapsulate specific actions their application is capable of, such as database interactions, API calls, or email operations. SK allows these functions to be orchestrated seamlessly across mixed programming language environments. The real power of SK lies in its AI-driven orchestration capabilities. Instead of meticulously choreographing the LLM interactions by hand, SK lets developers use natural language to state a desired outcome or task. The AI automatically determines how to combine the relevant semantic functions to achieve this goal, which significantly accelerates development and lowers the skill barrier for using LLMs. SK can benefit enterprises when building LLM applications by simplifying the application process, reducing the cost and complexity of prompt engineering, enabling in-context learning and reinforcement learning, and supporting multimodality and (continued) technical knowledge and resources. Unfortunately, no universal boilerplate code is available to develop an LLM orchestrator. Before proceeding with this project, consider the following factors: Customization—Tailoring the framework to meet your specific application and performance requirements Integration with existing systems—Seamlessly integrating the orchestrator with your existing infrastructure and workflows Control and visibility—Maintaining complete control over the LLM technology and accessing detailed insights into its operation Flexibility and scalability—Designing the framework to be flexible enough to accommodate future changes and scaling to meet growing demands If you want to create something entirely new, you need to understand generative AI, the different types of LLMs, how to train and fine-tune them, and how to use them for various tasks and domains. Additionally, you should know how to gather, process, and store data and knowledge that can help improve the quality and diversity of the generated outputs. To apply these concepts in real-world scenarios, you must be able to design and implement different generative strategies, such as prompt engineering and RAG. These strategies can help control the behavior and output of the LLMs. You must also ensure that the generative models and workflows are scalable, secure, and reliable. This can be achieved using cloud services, APIs, and UIs. Expertise in distributed systems, ML, and software engineering is also required. 29910.3 Orchestration layer multilanguage scenarios. SK provides a consistent and unified interface for different LLM providers, such as OpenAI, Azure OpenAI, and Hugging Face. Combining simplified LLM integration with AI-powered orchestration creates a powerful platform for enterprises to use to revolutionize their applications. Furthermore, SK makes it feasible to build highly tailored, intelligent customer support systems, implement more powerful and semantically nuanced search functionality, automate routine workflows, and potentially even aid developers with code generation and refactoring tasks. Additional details on SK can be found on their site at https://aka.ms/semantic-kernel. We can illustrate this using an example. Continuing with the pet theme from the previous chapters, we have some books about dogs, which range from general topics to more specific medical advice. These books are scanned and available as PDFs and contain confidential business data we want to use for a question–answer use case. These PDFs are complex documents that contain text, images, tables, and so forth. Given that we cannot use real-world internal information, these PDFs represent proprietary internal information for an enterprise that requires RAG to handle. Suppose we want to do question–answer use cases with the PDFs we have; let’s see how that’s possible. The first step is to use SK to install the SDK (or the package), which is not supported via conda and will require pip instead. Also note there are breaking changes with some of the SDKs, and we will want to pin the SK SDK to version 1.2.0. You can install this specific version using pip install semantic-kernel==1.2.0. After installing the SDK, to get started with SK at a high level, we need to follow these steps: Create an SK instance, and register the AI services you want to use, such as OpenAI, Azure OpenAI, or Hugging Face. Create semantic functions that are prompts with input parameters. These functions can call your existing code or other semantic functions. Call the semantic functions with the appropriate arguments, and await the results. The results will be the output of the AI model after executing the prompt. Optionally, we can create a planner to orchestrate multiple semantic functions based on the user input. SK EXAMPLE Here is an example of implementing this using the SK. As we saw earlier, SK is the core component that enables the processing and understanding of natural language text. It’s a framework that provides a unified interface for various AI services and memory stores. Our example is a simple question-answering system that uses the OpenAI API to generate embeddings for a collection of PDF documents. Then, we use those embeddings to find documents relevant to a user’s query. In our example, it is used for Creating embeddings—SK provides a simple interface for calling the OpenAI service to generate embeddings for the text extracted from PDF documents. As we 300 CHAPTER 10 Application architecture for generative AI apps know, these embeddings are numerical representations of the text that capture its semantic meaning. Storing and retrieving information—We use a vector database (Chroma in our example) to store the text and corresponding embeddings. SK calls these persistent data stores “memory” and, depending on the provider, has methods for querying the stored information based on semantic similarity. As we know, this is used to find documents relevant to a user’s query. Text completion—We also use SK to register an OpenAI text completion service, which is used to generate completions for a given piece of text. NOTE We need to specifically use Chroma version 0.4.15, as at the moment, there is an incompatibility with version 0.4.16 and higher with SK that hasn’t been fixed. To do this, we can use one of the following commands depending on whether we are using conda or pip: conda install chromadb=0.4.15 or pip install chromadb==0.4.15. Listing 10.1 shows this simple application processing a collection of PDF documents, extracting their text, and then using the OpenAI API to generate embeddings for each document. These embeddings are then stored in a vector database, which can be queried to find documents that are semantically similar to a given input. The load_pdfs function reads PDF files from a specified directory. It uses the PyPDF2 library to open each PDF, extract the text from each page, and return a collection of those pages. import asyncio from PyPDF2 import PdfReader import semantic_kernel as sk from semantic_kernel.connectors.ai.open_ai import ➥(AzureChatCompletion,AzureTextEmbedding) from semantic_kernel.memory.semantic_text_memory ➥import SemanticTextMemory from semantic_kernel.core_plugins.text_memory_plugin ➥import TextMemoryPlugin from semantic_kernel.connectors.memory.chroma import ➥ChromaMemoryStore # Load environment variables AOAI_KEY = os.getenv("AOAI_KEY") AOAI_ENDPOINT = os.getenv("AOAI_ENDPOINT") AOAI_MODEL = "gpt-35-turbo" AOAI_EMBEDDINGS = "text-embedding-ada-002" API_VERSION = '2023-09-15-preview' PERSIST_DIR = os.getenv("PERSIST_DIR") VECTOR_DB = os.getenv("VECTOR_DB") Listing 10.1 Q&A over my PDFs: Extracting text from PDFs 30110.3 Orchestration layer DOG_BOOKS = "./data/dog_books" DEBUG = False VECTOR_DB = "dog_books" PERSIST_DIR = "./storage" ALWAYS_CREATE_VECTOR_DB = False # Load PDFs and extract text def load_pdfs(): docs = [] total_docs = 0 total_pages = 0 filenames = [filename for filename in ➥os.listdir(DOG_BOOKS) if filename.endswith(".pdf")] with tqdm(total=len(filenames), desc="Processing PDFs") ➥as pbar_outer: for filename in filenames: pdf_path = os.path.join(DOG_BOOKS, filename) with open(pdf_path, "rb") as file: pdf = PdfReader(file, strict=False) j = 0 total_docs += 1 with tqdm(total=len(pdf.pages), ➥desc="Loading Pages") as pbar_inner: for page in pdf.pages: total_pages += 1 j += 1 docs.append(page.extract_text()) pbar_inner.update() pbar_outer.update() print(f"Processed {total_docs} PDFs with {total_pages} pages.") return docs After we have extracted the text from the pages, we use the populate_db() function to generate embeddings and store them in Chroma, a vector database. This function takes an SK object and goes through all the pages of the PDF. Each page saves the document’s text using the SK’s memory store. When the save_information() function is called, it automatically creates embedding to store in the vector database, as shown in the next listing. If there is already a Chroma vector database, we use that instead of making a new one. # Populate the DB with the PDFs async def populate_db(memory: SemanticTextMemory, docs) -> None: for i, doc in enumerate(tqdm_asyncio.tqdm(docs, desc="Populating DB")): if doc: #Check if doc is not empty try: await memory.save_information(VECTOR_DB,id=str(i),text=doc) except Exception as e: print(f"Failed to save information for doc {i}: {e}") continue # Skip to the next iteration Listing 10.2 Q&A over my PDFs: Using SK and populating vector database 302 CHAPTER 10 Application architecture for generative AI apps # Load the vector DB async def load_vector_db(memory: SemanticTextMemory, ➥vector_db_name: str) -> None: if not ALWAYS_CREATE_VECTOR_DB: collections = await memory.get_collections() if vector_db_name in collections: print(f" Vector DB {vector_db_name} exists in the ➥collections. We will reuse this.") return print(f" Vector DB {vector_db_name} does not exist in the collections.") print("Reading the pdfs...") pdf_docs = load_pdfs() print("Total PDFs loaded: ", len(pdf_docs)) print("Creating embeddings and vector db of the PDFs...") # This may take some time as we call embedding API for each row await populate_db(memory, pdf_docs) The program’s entry point is the main() function, as shown in listing 10.3. It sets up the SK with the OpenAI text completion and embedding services, registers a memory store, and loads the vector database. Then, it enters a loop where it prompts the user for a question, queries the memory store for relevant documents, and prints the text of the most relevant document. async def main(): # Setup Semantic Kernel kernel = sk.Kernel() kernel.add_service(AzureChatCompletion( service_id="chat_completion", deployment_name=AOAI_MODEL, endpoint=AOAI_ENDPOINT, api_key=AOAI_KEY, api_version=API_VERSION)) kernel.add_service(AzureTextEmbedding( service_id="text_embedding", deployment_name=AOAI_EMBEDDINGS, endpoint=AOAI_ENDPOINT, api_key=AOAI_KEY)) # Specify the type of memory to attach to SK. # Here we will use Chroma as it is easy to run it locally # You can specify location of Chroma DB files. store = ChromaMemoryStore(persist_directory=PERSIST_DIR) memory = SemanticTextMemory(storage=store, ➥embeddings_generator = kernel.get_service("text_embedding")) kernel.add_plugin(TextMemoryPlugin(memory), "TextMemoryPluginACDB") await load_vector_db(memory, VECTOR_DB) Listing 10.3 Q&A over my PDFs: SK using Chroma 30310.3 Orchestration layer while True: prompt = check_prompt(input('Ask a question against ➥the PDF (type "quit" to exit):')) # Query the memory for most relevant match using # search_async specifying relevance score and # "limit" of number of closest documents result = await memory.search(collection=VECTOR_DB, ➥limit=3, min_relevance_score=0.7, query=prompt) if result: print(result[0].text) else: print("No matches found.") print("-" * 80) if __name__ == "__main__": asyncio.run(main()) In our example, we use Chroma as the vector database. This is one of the many options available when using SK. We can get more details on the list of supported vector databases at https://mng.bz/YVgQ. It is also important to note that support between C# and Python is not at parity; some vector databases are supported across both, but some are only supported in one language. The SK is the central component for processing and understanding text. It provides a unified interface for various AI services and memory stores, simplifying the process of building complex NLP applications. Now let’s switch gears and see the same example using LangChain. LANGCHAIN LangChain offers a sophisticated framework designed to streamline the integration of LLMs into enterprise applications. This framework abstracts the complexities of interfacing with LLMs, allowing developers to incorporate advanced NLP capabilities without deep expertise in the field. Its library of modular components enables the construction of customized NLP solutions easily, facilitating a more efficient development process. LangChain’s main benefit is its ability to work with different LLMs and other natural language AI services. This feature allows enterprises to select the best tools for their particular needs, avoiding the drawbacks of being tied to one vendor. The framework boosts efficiency by providing easier interfaces and ready-made components for quick deployment and supports scalability, thus enabling projects to expand smoothly from testing stages to full-fledged applications. Additionally, LangChain helps to lower costs by minimizing the amount of specialized development and simplifying interactions with LLMs. Enterprises also gain from the strong community and support of the ecosystem around LangChain, which gives access to documentation, best practices, and cooperative problem-solving resources. This comprehensive approach makes LangChain an attractive option for businesses 310 CHAPTER 10 Application architecture for generative AI apps For data to be useful in informing LLM outcomes, it must first undergo a rigorous cleansing and standardization process to ensure its quality. The architectural blueprint should include these preprocessing activities, such as deduplication, normalization, and error rectification. Integrated data quality tools should automate these tasks, providing LLMs with superior datasets. Data handling requires strict access controls for proper security and compliance, which is vital when working with sensitive information and following regulations. Data interaction needs strong authentication and authorization protocols. Data governance frameworks should specify access rights; furthermore, encryption should protect data at rest and in motion. Frequent compliance assessments are crucial for ensuring data quality and privacy. Following GDPR, HIPAA, or CCPA regulations is also important for ethical and lawful processing of personal data. A plugin enabling the integration into source systems is not a one-time static component of the architecture—it changes and adapts constantly. As businesses use or improve their new SoRs, the architecture must be built to allow simple integration or movement of data sources. For this, a flexible approach to integration is required, where new data sources can be connected with little change to the current system. The architecture should be designed to support different data formats and protocols. This ensures that data flows seamlessly from various systems to the LLM. To achieve this, custom APIs may need to be developed, middleware may have to be used for data transformation, and ETL processes capable of handling large volumes of data may have to be implemented. The data pipeline infrastructure for generative AI is complex and requires careful planning to handle the intricacies of enterprise-grade data landscapes. These will build on existing ETL and data warehousing investments but must factor in the new data types of embeddings. By strategically using a combination of tools for data ingestion, processing, storage, orchestration, and ML, enterprises can build powerful pipelines that provide their generative AI applications with a consistent flow of quality data. 10.4.2 Embeddings and vector management In earlier chapters of the book, we discussed the crucial role of model embeddings and representations. This is the stage where the complexity of language is distilled into machine-interpretable formats, specifically mathematical vectors. Text is transformed by embedding techniques and advanced feature extraction forms that result in a vector space representation of text. These vectors are not arbitrary; they encapsulate the semantic essence of words, phrases, or entire documents, mapping information into a compressed, information-rich, lower-dimensional space. OpenAI Codex is a prime example of this process. It can comprehend and generate human-readable code, making it a powerful tool for embedding programming and natural languages. This is a significant advantage for code generation and automation tasks. In contrast, Hugging Face provides an extensive suite of pretrained models that are finely tuned for diverse languages and tasks. They can adeptly handle embeddings ranging from brief sentences to intricate documents. 31110.4 Grounding layer These models distinguish themselves by their ability to grasp contextual word relationships beyond basic dictionary meanings. By considering the words in their vicinity, the generated embeddings provide a nuanced reflection of the word usage and connotations within specific contexts. This feature is essential for generative AI applications that aim to emulate human-like text production. It fosters outcomes that are not only coherent and context-aware but also semantically profound. As we saw in earlier chapters on RAG, various libraries are available for chunking data, and some offer auto-chunking capabilities. One such library, called Unstructured (https://github.com/Unstructured-IO/unstructured), provides open source libraries and APIs that can create customized preprocessing pipelines for labeling, training, or production ML pipelines. The library includes modular functions and connectors that form a cohesive system, which makes it easy to ingest, preprocess, and adapt data to different platforms. It is also efficient at transforming unstructured data into structured outputs. An alternative solution is using LangChain and SK, which we saw earlier. These libraries support common chunking techniques for fixed size, variable size, or a combination of both. In addition, you can specify an overlap percentage to duplicate a small amount of content in each chunk, which helps preserve context. After transforming vectors, it is crucial to manage them properly. Vector databases specially designed to store indexes and retrieve high-dimensional vector data are available. Some such databases include Redis, Azure Cosmos DB, Pinecone, and Weaviate, to name a few. These databases help with quick searches within large embedding spaces, making it easy to identify similar vectors instantly. For instance, a generative AI system can use a vector database to match a user’s query with the most semantically related questions and answers and achieve this in a fraction of a second. Vector databases feature sophisticated indexing algorithms engineered to deftly traverse high-dimensional terrains without falling prey to the “curse of dimensionality” [5]. This attribute renders them exceptionally valuable for applications such as recommendation engines, semantic search platforms, and personalized content curation, where pinpointing relevant content quickly is critical. Vector databases offer more than just speed; they also provide accuracy and relevance. Combining these databases allows AI models to respond quickly and precisely to user inquiries based on their learned context. Proper index management is crucial, including tasks such as index creation, update triggers, refresh rates, complex data types, and operational factors (e.g., index size, schema design, and underlying compute services). Cloud-based solutions such as Azure AI Search and Pinecone can efficiently manage these demands in a production environment. The process of transforming textual data into a format that AI can handle has two stages: embedding and vector database management. This conversion is essential for generative AI’s intelligence, enabling it to understand and engage with the world meaningfully and in a scalable manner. Therefore, carefully choosing embedding techniques and vector databases is a technical necessity and a key factor in the success of generative AI applications. When choosing LLMs, related vector storage and 312 CHAPTER 10 Application architecture for generative AI apps retrieval engines, and embedding models, enterprises must consider the data size, origin, change rate, and scalability needs. 10.5 Model layer The model layer is the foundation of AI cognitive capabilities. It involves a set of models, including foundational LLMs that provide general intelligence, fine-tuned LLMs specialized for specific tasks or domains, model catalogs hosting and managing access to various models, and SLMs that offer lightweight, agile alternatives for certain applications. The significance of this layer lies in its design, as it forms the core processing units of the GenAI app stack. It allows a scalable and flexible approach to AI deployment and can efficiently address various tasks by differentiating between foundational, fine-tuned, and small models. This ensures that the architecture can cater to diverse use cases, optimize resource allocation, and maintain high performance across different scenarios. 10.5.1 Model ensemble architecture Generative AI employs a model ensemble, which combines multiple ML models to enhance performance and reliability. This approach takes advantage of the individual strengths of each model, minimizing their weaknesses. For example, one model may be great at generating technical content, while another may be better at creative storytelling. By assembling these models, an application can better cater to a wider range of user requests with greater accuracy. To create an effective model ensemble for generative AI, the architecture should include Model selection—Criteria for choosing which models to include in the ensemble, often based on their performance, the diversity of training data, or their area of specialization. Small language models SLMs such as Phi-3 and Orca 2 are designed to offer advanced language processing capabilities with fewer parameters than larger models. Both models are part of a broader initiative to make powerful language processing tools more accessible and efficient, enabling more extensive research and application possibilities. They represent a significant step in the evolution of AI language models, balancing capability with computational efficiency. Phi-3, Phi-2, and Orca 2 are smaller-scale language models developed by Microsoft, offering advanced language processing with fewer parameters. Phi-3, which is a successor to Phi-2, is a family of models in various sizes (mini, 3.8B; small, 7B; medium, 14B parameters). Phi-2, with 2.7 billion parameters, is efficient and matches larger models in performance, while Orca 2, available in 7and 13-billion-parameter versions, excels in reasoning tasks and can outperform much larger models. Both are designed for accessibility and computational efficiency, enabling broader research and application in AI language processing. 31310.5 Model layer Routing logic—Routing logic is the mechanism for determining which model to use for a given input or how to combine outputs from multiple models. API integration—APIs are the main conduits through which applications interact with LLMs. API integration becomes complex when dealing with an ensemble of models as interactions with multiple endpoints must be managed. The architecture should consider API integration of throttling and rate limits, error handling, and caching responses. Scalability and redundancy—Scalable design accommodates growing user bases and spikes in demand. Load balancing and the use of API gateways can help distribute traffic effectively. Redundancy is equally critical; thus, having multiple regions for model deployments ensures the application remains functional. Queuing and stream processing—Queuing and stream processing handle asynchronous tasks and manage workloads; message queues and stream processing services can be utilized, which ensures that the system is not overwhelmed during peak times and that tasks are processed in an orderly way. Figure 10.5 is an example of implementing Phi-2 as a classifier. We use Phi-2, which runs locally and fast, to identify the user’s intent when asking a question. Continuing with the topic of pets and dogs, we asked Phi-2 the intent of the question and whether it had anything to do with dogs. If it was irrelevant to the current topic (i.e., dogs), we asked GPT-4 to answer. Figure 10.5 Classifier using multiple models Listing 10.6 shows an example of implementing a simple classifier using a lightweight model and then, based on the question’s intent, figuring out which model to call. Here, we use Phi-2, a research SML from Microsoft, as a classifier to determine whether a question is related to dogs. The Phi-2 model is a transformer-based model, trained to understand and generate human-like text. It is used here as a first-pass filter to determine the question’s intent. User question Question about dogs? Classification Formulate prompt Phi 2 — Intent classifier (SLM) Response to user Azure OpenAI GPT-4 model Yes Local inference 314 CHAPTER 10 Application architecture for generative AI apps The function check_dog_question() takes a question as input and constructs a prompt to ask the Phi-2 model whether there’s anything about dogs in the question. If Phi-2 determines that the question is about dogs, the function returns True. This could trigger a more expensive GPT-4 model to generate a more detailed response. If the question is not about dogs, the function returns False, and the more expensive model would not have to be used. We need to ensure that the following packages are installed before running this code: pip install transformers==4.42.4 torch= =2.3.1. import torch from transformers import AutoModelForCausalLM, AutoTokenizer import openai ... model = AutoModelForCausalLM.from_pretrained("microsoft/phi-2", torch_dtype="auto", trust_remote_code=True) tokenizer = AutoTokenizer.from_pretrained("microsoft/phi-2", trust_remote_code=True) def check_dog_question(question): prompt = f"Instruct: Is there anything about dogs in the ➥question below? If yes, answer with 'yes' else ➥'no'.\nQuestion:{question}\nOutput: " inputs = tokenizer(prompt, return_tensors="pt", return_attention_mask=False, add_special_tokens=False) outputs = model.generate(**inputs, max_length=500, pad_token_id=tokenizer.eos_token_id) text = tokenizer.batch_decode(outputs)[0] regex = "^Output: Yes$" match = re.search(regex, text, re.MULTILINE) if match: return True return False def handle_dog_question(question): print( "This is a response from RAG and GPT4") # Call OpenAI's GPT-4 to answer the question openai.api_key = "YOUR_API_KEY" response = openai.Completion.create( … ) return response Listing 10.6 Using Phi-2 as an intent classifier 31510.5 Model layer if __name__=="__main__": # Loop until the user enters "quit" while True: # Take user input user_prompt = input( "What is your question (or type 'quit' to exit):") if check_dog_question(user_prompt): print(handle_dog_question(user_prompt)) else: print("You did not ask about dogs") The approach employs a small model, such as Phi-2, with much less capability for more efficient use of resources, as the more expensive GPT-4 model is used only when necessary. This approach can just as easily be expanded to use more than one model. This toy example could be better if we used a more powerful LLM, such as a smaller GPT-3 model. Figure 10.6 shows another example of using a fine-tuned GPT-3 as a classifier to help understand the user’s goal. This is for an enterprise chatbot that can answer questions on both structured and unstructured data. It can answer questions about Microsoft’s surface devices based on the user’s persona. There is fictitious sales information in a SQL database that a salesperson can chat with, and there is unstructured data that can answer technical support questions. Figure 10.6 Enterprise Q&A bot—High-level overview The bot uses a RAG pattern and can answer questions using information from both structured and unstructured systems based on the user’s intention. The structured data has sales information (with fake data), and the unstructured data is a crawl of different forums and official sites related to Surface devices. Listing 10.7 presents a highlevel view of the architecture. User question Question [Intent + Domain] Classification Formulate prompt Azure OpenAI service — Fine-tuned GPT intent classifier Response to user Azure OpenAI ChatGPT model Structured and unstructured knowledge bases [Document text truncated for crawler view.]