Share E-Book

The Enterprise Data Catalog Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture (Early Release) (Ole Olesen-Bagneux) (z-library.sk, 1lib.sk, z-lib.sk)

Author Ole Olesen-Bagneux

ai
Language English

It's a new day in search. Before ChatGPT, combing the web was simple—powerful search engines dominated for 25 years. That changed with conversational search powered by chatbots. Data catalogs, the search engines for your company's data, have evolved as well. The Enterprise Data Catalog explores how AI is transforming enterprise-wide data search. In this second edition, you'll explore how the role of the data catalog has changed in the age of AI. Data catalogs no longer serve as tools to find and use data—they now deliver essential metadata for AI projects. Author Ole Olesen-Bagneux explains how metadata organized as enterprise ontologies, delivered through knowledge graphs, provides the context required by large language models, Model Context Protocol, and Agent2Agent Protocol. By drawing on data management and library and information science, the book shows why information science methodology is critical to successful catalog implementations. • Understand the role of the enterprise data catalog in discovery, governance, and AI • Organize data and sources using metadata, domains, and enterprise ontologies • Search and browse data across domains, lineage, and graphs • Leverage data catalog knowledge graphs for AI use cases • Apply data catalogs to data products and data contracts

Format PDF
Size 2.5 MB
2
Views
0
Downloads
0.00
Total Donations
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
(This page has no text content)
Page 3
(This page has no text content)
Page 4
(This page has no text content)
Page 5
With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write— so you can take advantage of these technologies long before the official release of these titles. Ole Olesen-Bagneux The Enterprise Data Catalog Scale AI with Metadata Using LLMs, MCP, and Agentic Architecture SECOND EDITION
Page 6
979-8-341-67290-1 [LSI] The Enterprise Data Catalog by Ole Olesen-Bagneux Copyright © 2027 O’Reilly Media, Inc. All rights reserved. Published by O’Reilly Media, Inc., 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (https://oreilly.com). For more information, contact our corporate/institu‐ tional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Aaron Black Development Editor: Sara Hunter Production Editor: Katherine Tozer Interior Designer: David Futato Interior Illustrator: Kate Dullea February 2023: First Edition June 2027: Second Edition Revision History for the Early Release 2026-02-18: First Release 2026-03-20: Second Release 2026-06-05: Third Release See https://oreilly.com/catalog/errata.csp?isbn=9798341672949 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. e Enterprise Data Catalog, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the author and do not represent the publisher’s views. While the publisher and the author have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the author disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. This work is part of a collaboration between O’Reilly and Actian. See our statement of editorial independ‐ ence.
Page 7
Table of Contents Brief Table of Contents (Not Yet Final). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . vii Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix 1. Introduction to Data Catalogs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 The AI Data Catalog and as Source for AI 18 The Core Functionality of a Data Catalog 19 Create an Overview of the Data in the IT Landscape 20 You Will Always Be Able to See More Metadata than Data 21 Organize Data 21 Enable Search of Company Data 25 Data Discovery 29 The Data Discovery Team 31 Data Catalog Ownership 32 End-User Roles and Responsibilities 33 Summary 35 2. Organize Data: Design a Robust Architecture for Search. . . . . . . . . . . . . . . . . . . . . . . . . . 37 Using AI to Organize Data 38 Organizing Domains in the Data Catalog 39 Domain Architecture in a Data Catalog 39 Understanding Domains 42 Processes and Capabilities 45 Data Sources 49 Organizing Data in the Domains 51 Metadata for Data 51 Knowledge Graph Powered Data Catalogs 56 Data Assets and Data Products 57 v
Page 8
Metadata Quality 60 Classification 64 Summary 67 3. Search For Data: Concepts, Features, Mechanics. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 Concepts 70 Searching in Data Versus Searching for Data 70 Information Needs: Search Like Librarians—Not Like Data Scientists 75 Serendipity 77 Promptism 78 Features: Search Features in a Data Catalog 79 Simple Search 81 Browsing 84 Complex Search 87 Conversational Search 89 Mechanics: The Mathematics Behind Search 90 Recall and Precision 91 Zipf ’s Law 94 Summary 96 4. Search For Data Patterns. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 Search Pattern Overview 99 Keyword Search and Conversational AI Search 100 Basic Simple Search 101 Detailed Simple Search 102 Flexible Simple Search 104 Range Search 105 Block Search 106 Statement Search 110 Browsing Patterns 111 Glossary Browsing 111 Domain Browsing 112 Lineage Browsing 113 Graph Browsing 113 Searching a Graph-Based Data Catalog 115 Summary 116 vi | Table of Contents
Page 9
Brief Table of Contents (Not Yet Final) Preface (available) Chapter 1: Introduction to Data Catalogs (available) Chapter 2: Organize Data: Design a Robust Architecture for Search (available) Chapter 3: Search Data: Concepts, Features, Mechanics, Patterns (available) Chapter 4: Search for Data Patterns (available) Chapter 5 Access and Observe Data (unavailable) Chapter 6 Empower End Users and Engage Stakeholders (unavailable) Chapter 7 Data Domains (unavailable) Chapter 8 Data Architecture, Providers, and Consumers (unavailable) Chapter 9 Data Products and Data Contracts (unavailable) Chapter 10 Manage Data: Improve Lifecycle Management (unavailable) Chapter 11 e Data Catalog Is Now a Source in Itself (unavailable) Chapter 12 e LLM + KG Pattern (unavailable) Chapter 13 Standards and AI (unavailable) vii
Page 10
(This page has no text content)
Page 11
1 Monday Morning Data Chat was a podcast by Joe Reis and Matt Housley, I was interviewed in the episode The Future of Data Catalogs, Matt asks the question 52:10 Preface A Note for Early Release Readers With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. This will be the Preface of the final book. If you’d like to be actively involved in reviewing and commenting on this draft, please reach out to the editor at shunter@oreilly.com. It’s A New Day in Search “So what’s your take on ChatGPT? How does it change data catalogs? Do you have an opinion about that?” Matt Housley was putting me on the spot. I was invited on the podcast Monday Morning Data Chat1 and we were discussing my soon to be published book e Enterprise Data Catalog, the first edition of this book that you are now reading in the second edition. I wasn’t properly mic’ed up, I was not yet used to being on tech podcasts and painfully aware of my clunky sound. I pulled myself together and managed to answer. I said that my book ended with a future vision for data catalogs, and that if I was to go into more depth about that future vision then, obviously I would look into the potential of ChatGPT and AI in general. ix
Page 12
2 NPR, Feb 7th, 2023: Microsoft revamps Bing search engine to use artificial intelligence It was a paradoxical moment, really. As I was publishing the first edition of e Enter‐ prise Data Catalog, in early 2023, the core message of the book was that data catalogs are like search engines, just for data in companies. That’s why data catalogs need to be built on knowledge graphs, just like search engines are. Don’t worry, we will get to knowledge graphs, they play a major role in this book! And yet, search engines themselves, the real ones, for the web, were unexpectedly challenged, right there in early 2023, as I published e Enterprise Data Catalog. For 25 years, the biggest business on the web - search - had been completely stable. Not technologically, of course it had evolved, but in terms of how you searched the web, as an end user. You had one, simple search bar, that would provide the best, the freshest, the most relevant search hits to you, with the blink of an eye. For 25 years, we had been using search engines as the most natural extension of our mind to search for everything from our absolute basic needs to the most complex types of curiosity the human brain can produce. And then, all of a sudden, it was over. In November 2022, ChatGPT 3.5 was released by OpenAI. Suddenly, we asked ourselves if the way we search was in fact obsolete. Would we talk with the web instead? Do conversational search through a chatbot? In those very early months of 2023 Microsoft CEO Satya Nadella said: “It’s a new day for search” when Microsoft joined forces with OpenAI. Nadella, Microsoft, was chal‐ lenging Google: Microsoft was powering their search engine Bing with conversational search in a chatbot supported by OpenAI.2 The race for winning a disruption for search, as a business, was on - it was a spectacular turn of events after more than two decades of complete domination by Google. Nothing captured the zeitgeist better than the cover of the Economist 11th Feb 2023, as depicted in figure x.1. The battle for search was on. x | Preface
Page 13
Figure P-1. e Economist nailed what AI did to search engines And here I was, arguing that data catalogs were like search engines, just for the data in your organization. And so, back to Matt’s question: How did this new technological push forward in AI change data catalogs? Preface | xi
Page 14
At the time I was asked the question, it was too early to provide an answer. No one knew. But something was clearly going to change. I saw things that I knew from my academic background would never fly, but I also saw interesting experiments. The first edition of my book was very well received, I was lucky to have thousands of readers all over the planet. The book was praised in reviews and featured in many podcasts, and I got to travel the world explaining my ideas at big conferences, events and for large corporations - the book even got translated. Its success was its message. Data catalogs are engines to search for data, so that you could later search in data, directly in databases. Static, preconceived metamodels in data catalogs led to catastrophic implementations, sky-rocketing costs and no ROI - because that is basically not how a search engine should work, not on the web, not in a company. As a PhD and an aspiring professor, I had taught knowledge organization, informa‐ tion retrieval and adjacent topics in my Library- and Information Science classes at the University of Copenhagen in Denmark, where I live. I went into industry, but I never forgot what I learned and taught. And it was on that basis, that I argue, that data catalogs, essentially, are like search engines. And I wrote e Enterprise Data Catalog from that perspective. And then came AI. All of a sudden there was a new dimension to data cataloging that I had not covered in my book. Why I Wrote This Book Therefore, it is time for a second edition of e Enterprise Data Catalog. All the insights from the first edition are kept - it all holds true still. The AI perspective has been added on top, and it falls into two parts: • Part I: How AI augments data catalogs • Part II: How data catalogs (ontologies) are a source for AI You will notice that these two parts are also the parts of this book - because we can implement and use data catalogs at scale with AI, and we can use the very structure of the data catalog, the metadata, as a source in itself, to fuel AI. Basically, because of AI, it’s also a new day in search for data catalogs. We are now beginning to understand what that day is about, and that is what we will unfold in the following pages. This book is about how AI augments the core message of my book, namely; how you organize data, denes how you can search it. And this message also has a new dimension, since the data catalog is in itself a source. Finally, the data mesh movement was at its peak when I published the first edition of e Enterprise Data Catalog. At the time of publication, it was difficult to say what xii | Preface
Page 15
would remain from the movement, after the buzz would fade. Now, that has become clear. No one talks about data mesh anymore, but two of the components in the data mesh complex stood the test of time: data products and data contracts. Therefore, this second edition covers data products and data contracts, because they are absolute key components in data catalogs. Who Should Read This Book This book is for everyone working in data, meaning • Data engineering • Data analysis • Data science • Data management • Data governance These groups of employees all work with data and they all need to know where data is, who owns data, the quality of data, how data moves, where data ends up, and who uses it. The data engineers facilitate the storage, transformation, and movement of data, and they can monitor that in the data catalog. The data analysts and scientists use the data catalog to search for interesting data - and then, once found, they will search in those data sources. Data managers and data governors use the catalog as a strategic tool to ensure that data governance policies are successfully implemented. On top of that, a wealth of groups can benefit from using a data catalog, e.g. • Architects • Information security • Data protection officers Architects can use a data catalog for effective data migration projects, information security and data protection to ensure an empirical validation of their anticipation of what data the organization has. Navigating This Book This book has two parts, Part I: e AI augmented Data Catalog, and Part II: Data Catalogs as a resource for AI. Part I: e AI Augmented Data Catalog is all about how data catalogs work and how they have improved since the first edition - mainly thanks to AI. The chapters intro‐ duce data catalogs, explains how data is organized in data catalogs and how search is Preface | xiii
Page 16
performed, furthermore how data is accessed and observed. All these aspects are improved, made easier and faster, thanks to AI. Also, we will dive deeper into data domains, data architectures, and especially focus on data products and data contracts, as these matured significantly since the first edition. Part II: Data catalogs as a resource for AI uncovers a new perspective: Data catalogs are now becoming sources themselves, not only a catalog of sources. This new role for metadata is discussed in the light of the combination of knowledge graphs and large language models. Furthermore, the emerging standards of effective AI are explained, such as Model Context Protocol and Agent2Agent. This second ends with a future vision for data catalogs, and this is substantially updated since the first edi‐ tion. Preface for the First Edition “This simply can’t be all there is to a data catalog. What does it really do?” About five years ago, I sat alone in the office among 20 empty desks. My company had shut off the air-conditioning to go green, so I was uncomfortably warm on top of being perplexed by the bunch of white papers, both printed and on my laptop, that were sitting in front of me. The papers explained a new technology called a data cata‐ log. As an enterprise architect, I had been asked to implement a data catalog for our company. But first, I had to understand it. The papers I was looking at described cool, advanced features: column-based data lineage, graph visualizations of ontologies, and workflows to access virtualized data. Useful. Mesmerizing, really. But what was the overall point of a data catalog? I was sweating, physically and mentally, trying to draw upon my experiences to figure out the potential of this new technology. I have a BA, MA, and PhD in library and information science (LIS). I have taught LIS in university courses and been to conferences all over the world. I’ve seen a lot of things in this field, both good and bad. During my first job in pharma, senior man‐ agement regularly called me late at night because inspections from the authorities were going haywire. The inspectors were asking them a multitude of questions: What was the temperature of this tube, in that machine, in June 1992? Where is the proof that the fermentation tank was cleaned according to the standard operating proce‐ dure (SOP)? When the data managers searched and couldn’t answer, they called my team—the Records and Information Management team. We employed our searching superpowers to find the information they needed. We were adept with our queries, used intuition and creativity to plan our moves, and drew on our knowledge to guide our search. We were able to do this because we had one guiding principle: How you organize data denes how you can search it. Because we knew how the data was organized, we knew ways to begin searching, modifying xiv | Preface
Page 17
how we searched, broadening, changing focus, excluding hits, and finding the infor‐ mation we needed. Sometimes this was easy, and sometimes it was hard, but we would get there every time. This guiding principle has followed me throughout my career. I have cataloged furni‐ ture, weapons, human tissue, a lot of paper, and massive amounts of data. I know how to structure and operate a physical card catalog in a library. I know records and infor‐ mation management systems with both physical and digital storage. We both stored and cataloged data on premises, and then, later, in the cloud. Throughout everything I experienced, I saw that if you have a poorly organized data landscape, searching for the information you need will be a terrible experience. You will have to guess where to search and what to query. If your data is logically and systematically organized, however, you will know exactly where to look and what to query. It will be a much better experience. The idea that how you organize data defines how you search for it is reflected in our web habits as well. We never really think about how we search it anymore; it’s so intu‐ itive. At work, within our company’s IT landscape, it can be a different story. We search in vain—companies hardly know their own data, let alone how it is processed. Data is undiscoverable and unmanageable. If only we had an enterprise search engine… On that hot summer day, alone in the empty office, surrounded by physical papers and dozens of open PDFs on my laptop, it suddenly hit me. “This data catalog has the potential to become a search engine for companies! We are finally getting an engine that will be able to do for companies what search engines have done for the web. The data catalog is a search engine!” That realization led to another a few years down the road. All of the papers I read that fateful day, along with all of the documentation that followed it, focused on explain‐ ing the complex features that are in data catalogs. They did not explain the data cata‐ log itself and how it could revolutionize how we organize and search data. Nowhere has anyone talked about the future of the data catalog as an enterprise search engine. That realization has brought us to the book you are reading today. Although I had the epiphany about the potential of a data catalog and it was crystal clear in my head, I was then faced with the battle of explaining the features to the important stakeholders of my company. Although I knew they would see the benefits of this tool if they only took the time to understand it, they were simply not interes‐ ted, nor did they have time to study it. I had to come up with a way to reach them. I went back to the idea of using the data catalog as an enterprise search engine. So, I asked myself, “What are people searching for? What would a data scientist be search‐ ing for? A data protection officer? A chief information security officer?” Preface | xv
Page 18
I decided to build demonstrations of the most vital data catalog features into small stories about specific stakeholders. Each slide deck had one central picture: a mini‐ malist search bar with the company’s logo above it. I would explain the information need of a specific stakeholder, show the search in the search bar, reveal the result, then close with how the results could be used. In this way, I showed simple searches, complex searches, how to browse back and forth in the lineage of data, up and down in domains, and relationally in the graph that depicted our company. It had the same content as my previous demonstrations, but this time, it was explained from a stake‐ holder point of view: a specific person who was searching for something specific. And that worked. The stakeholders not only got interested, but they also got excited. They now wanted the data catalog, because they understood that this tool was not just a collection of fancy features for data geeks. No, this tool was something way more fundamental: the data catalog could help them search and find the data they were looking for. I explained that, implemented with care, a data catalog has relevance for many of the employees in a company. This approach worked for me and my colleagues, and I hope that it will work for you and yours as well. At the end of the day, we are all searching for something. And we search all the time. The only thing is, at work, it is very difficult to search for whatever we are trying to find. And we take that for granted, as something that we must just accept. I’m assuming you’re reading this book because you’re involved with planning to implement a data catalog, improve an existing one, sunset it, or simply trying to understand what kind of technology a data catalog is: what it does, how it should be used, and if it can help you in a certain way. You might be part of the offices of the legal counsel, chief data officer, data protection officer, or chief information officer. You might be a data engineer, data scientist, or data manager, or you might be part of the data governance team. If you are, then this book will help you understand what a data catalog is and how it will enable you to find exactly what you are searching for. However, you may also be a data catalog provider. In my book, I put forward a vision for the future of data catalogs, which you could benefit from when planning the future development of your data catalog technology. xvi | Preface
Page 19
CHAPTER 1 Introduction to Data Catalogs A Note for Early Release Readers With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. This will be the 1st chapter of the final book. If you’d like to be actively involved in reviewing and commenting on this draft, please reach out to the editor at shunter@oreilly.com. In this chapter, you’ll learn how a data catalog works, who uses them, and why. Before we dive into that, we will focus on why data catalogs have become relevant in a new way for Artificial Intelligence (AI). Pay close attention to this part, as it will ensure you strategic buy-in at the executive level, also for data governance and compliance, which is usually hard to get - as you are perhaps well aware. Then, we’ll go over the core functionalities of a data catalog and how it creates an overview of the data in your organization’s IT landscape. I’ll learn how the data can be organized in a data catalog, and how it makes searching for your data easy. Search is often underutilized and undervalued as part of a data catalog, which is a huge detri‐ ment to data catalogs. As such, we’ll talk about your data catalog as a search engine - and an AI assistant! - for your enterprise data that will unlock the potential for success. In this chapter, you’ll also learn about the benefits of a data catalog in an organiza‐ tion: a data catalog improves data discoverability, subsequently ensuring data gover‐ nance and enhancing data-driven innovation. Moreover, you’ll learn about how to set 17
Page 20
1 There will be references to relevant scientific articles documenting this, in Part II also. up a data discovery team and you’ll learn who the users of your data catalog are. I’ll wrap up this chapter by explaining the roles and responsibilities in the data catalog. OK, off we go. The AI Data Catalog and as Source for AI Before we jump into the nuts and bolts of what a data catalog is, let’s take a moment and focus on why data catalogs have become relevant both • with AI, and • for AI. With AI. Data catalogs are now substantially easier to implement, use and scale, thanks to many of the features in them being supported by AI. This will become clear already in this chapter, in the search examples below, and in we will go deeper into this in the other chapters of Part I. For AI. As you will also see in this chapter, data catalogs are now part of a bigger search infrastructure, via AI assistants. AI assistants not only search on data catalogs but in all sources connected to the AI assistant. This has created a remarkable shift: The data catalog is not only a tool leading to sources, it has become a source in itself. You might know the term AI assistant under the term chatbot. Think Claude, Mistral, OpenAI and others, only for dedicated enterprise use. Throughout this book, we use AI assistants. Furthermore, data catalogs - if they are built on knowledge graphs - are a rich source for AI. Graphs have proved to increase precision greatly when applying Large Lan‐ guage Models (LLM) for generative AI projects. Furthermore, agentic architectures are also executing tasks more effectively when supplemented by a knowledge graph.1 Overall, the combination of AI and knowledge graphs is promising and we will dis‐ cuss this in detail in Part II. You may think, If you are in data governance or data engineering: “Why does all this AI stuff matter to me?” I need my data catalog anyway. The answer is simple. It mat‐ ters to you because this evolution has made it significantly easier to get your enter‐ prise decision makers on board in implementing a data catalog, and you can also expect a smoother experience in rolling it out for new end users. 18 | Chapter 1: Introduction to Data Catalogs
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
Back to List