Share E-Book
Scan to open this page

Scan with your phone to open this page

AuthorPanos Alexopoulos

What value does semantic data modeling offer? As an information architect or data science professional, let’s say you have an abundance of the right data and the technology to extract business gold—but you still fail. The reason? Bad data semantics. In this practical and comprehensive field guide, author Panos Alexopoulos takes you on an eye-opening journey through semantic data modeling as applied in the real world. You’ll learn how to master this craft to increase the usability and value of your data and applications. You’ll also explore the pitfalls to avoid and dilemmas to overcome for building high-quality and valuable semantic representations of data. * Understand the fundamental concepts, phenomena, and processes related to semantic data modeling * Examine the quirks and challenges of semantic data modeling and learn how to effectively leverage the available frameworks and tools * Avoid mistakes and bad practices that can undermine your efforts to create good data models * Learn about model development dilemmas, including representation, expressiveness and content, development, and governance * Organize and execute semantic data initiatives in your organization, tackling technical, strategic, and organizational challenges

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical field guide for information architects, data scientists, and knowledge engineers who want to master semantic data modeling—not just the theory, but the real-world pitfalls and hard choices that determine whether your data models deliver value or quietly fail. 【Book Arc】 - **Opening (~0%–9%)**: Introduces the "semantic gap"—the disconnect between raw data and its meaning—and defines semantic data modeling as the craft of closing it. Sets up the book's structure: Part I covers basics, Part II examines pitfalls, Part III tackles dilemmas, and Part IV addresses organizational execution. - **Early (~9%–25%)**: Builds a common vocabulary for modeling elements across frameworks like RDF(S), OWL, SKOS, Description Logics, and database conceptual models. Covers entity types, relations, attributes, and lexical labels, using real examples like the ESCO skills taxonomy to show how these elements work in practice. - **Early (~25%–34%)**: Explores semantic and linguistic phenomena—instantiation, ambiguity, uncertainty, rigidity, unity, and semantic relations like subsumption and part-whole. Shows how these phenomena shape model quality and why getting them wrong creates downstream problems. - **Middle (~34%–47%)**: Shifts to quality dimensions, focusing on completeness and trustworthiness. Discusses gold standards, silver standards, and heuristics for measuring completeness, plus the subjective nature of trust in semantic models. - **Late (~47%–end)**: Moves from pitfalls to dilemmas—representation choices, expressivity vs. content trade-offs, and model evolution/governance. Concludes with forward-looking themes: avoiding tunnel vision, the semantic vs. nonsemantic framework debate, and symbolic knowledge vs. machine learning. 【Key Takeaways】 - **The semantic gap is the root cause of data failure** (Opening): Even with abundant data and powerful technology, projects fail when the meaning behind data is ambiguous, vague, or inconsistent. Semantic modeling is the discipline of making meaning explicit and machine-interpretable. - **Modeling elements are the building blocks, but terminology varies across frameworks** (Early): RDF(S) calls attributes "properties," OWL distinguishes "datatype properties," and SKOS has its own label system. Understanding these equivalences lets you analyze and compare models across different languages and standards. - **Lexical labels and synonyms are powerful but dangerous** (Early): They clarify meaning and help systems handle linguistic variety—like matching "software developer" and "software programmer" in job vacancy data. But false synonyms create errors, and over-including lexicalizations can bloat a model. - **Ambiguity must be resolved at modeling time, not left to chance** (Middle): When modeling statements like "John and Jane are married," you must decide whether it means a mutual relation or a shared attribute. Getting this wrong propagates confusion through every application that uses the model. - **Completeness is measurable but context-dependent** (Middle): Gold standards are rare, so use partial gold standards or silver standards to reveal incompleteness. A model can be complete for one use case (German stocks for a German investor) and incomplete for another (European stocks overview). - **Trustworthiness is social, not just technical** (Middle): A model can be less accurate yet more trusted, because trust involves perception and confidence, not just correctness. Building trust requires attention to governance and user communication, not just data quality. - **Semantic relations like subsumption and part-whole enable reasoning** (Middle): Properties such as symmetry (if John is Jane's brother, Jane is John's sister) let systems infer new knowledge. Recognizing when these properties truly hold—and when they don't—is critical for reasoning systems. 【Reading Tips】 - **Skim Part I if you're already familiar with RDF, OWL, and SKOS**—the terminology review is thorough but standard. Focus instead on the ESCO examples, which show real modeling decisions and their consequences. - **Deep-read the chapters on ambiguity and completeness** (Chapters 6 and 8 area). These are where the book earns its keep, with concrete guidelines for resolving meaning and measuring quality that you can apply immediately. - **Pay special attention to the pitfalls chapters** (Part II)—they're organized around common mistakes like presenting subjective knowledge as objective, which the ESCO "essential skills" example illustrates powerfully. - **Use the dilemmas chapters (Part III) as a decision framework**—when you face a choice between representation options or expressivity vs. content, return to these chapters to structure your thinking. - **Take away the quality dimensions as a checklist**—correctness, completeness, trustworthiness, and others can serve as a practical audit tool for your own models. 【Coverage Limits】 Excerpts cover the book's structure, core concepts, modeling elements, and quality dimensions well, but do not include detailed content from the later chapters on specific dilemmas (representation, expressivity, governance) or organizational execution. The final chapter's forward-looking themes are visible only as section titles.
Page 7
4 Why Develop and Use a Semantic Data Model? 7 Bad Semantic Modeling 8 Avoiding Pitfalls 10 Breaking Dilemmas 11 2. Semantic Modeling Elements. . . . . . . ....
View in text
Excerpt 2
, you attempt to distinguish between essential and optional skills, as ESCO does, then you should prepare for a lot of debate and disagreement. Just take a l...
View in text
Excerpt 3
, this relation is known as rdf:type while in the ANSI/NISO Standard it is indicated by the abbreviation BTI (standing for “Broader term (instance)”) or NTI...
View in text
Excerpt 4
d, it’s (knowingly) not completely accurate. Instead, it is assumed to have a reasonable level of quality that can be useful for detecting incom‐ plete aspec...
View in text
Excerpt 5
del’s ambiguity and, for that, we actually had to develop a whole new disambiguation module for the system. • Tackling any conflicting requirements that diff...
View in text
Excerpt 6
be interpreted by a human. Include multiple people in this. Alternatively, you could search the name within one or more corpora (or even Google it), and see...
View in text
Excerpt 7
along the dimension of the level of knowledge on a subject Ignoring Vagueness | 103 CHAPTER 7 Bad Semantics Words are wonderfully elastic. They can be mispro...
View in text
Excerpt 8
ned to you too, keep reading. Why We Get Bad Specifications The main mistakes I made during that first period of my work at Textkernel can be summarized as f...
View in text
Tags
AI categories
DataTechnologyProgramming
ISBN: 1492054275
Publisher: O'Reilly Media
Publish Year: 2020
Language: English
Pages: 329
File Format: PDF
File Size: 9.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…