Page
1
(This page has no text content)
Page
2
(This page has no text content)
Page
3
Data-Oriented Programming in Java
Page
4
welcome Thanks for purchasing the MEAP for Data-Oriented Programming in Java! This book is a distillation of everything I’ve learned about what effective development looks like in Java. It’s what’s left over after years of experimenting, getting things wrong (often catastrophically), and slowly having anything resembling “devotion to a single paradigm” beat out of me by the great humbling filter that is reality. Data Orientation is not some new paradigm here to beat up all other paradigms and take their lunch. If you like object orientation, functional programming, or any other paradigm, a little touch of data-orientation will make you better at those styles of programming! Data Orientation is about the data, not the specific tools. There are, of course, some patterns and approaches that naturally emerge when you focus on the data, but all of those can be applied readily to whatever paradigm is your preferred one. This is because data-orientation is born from a very simple idea, and one that people have been rediscovering over and over again since the dawn of computing: “representation is essence of programming”. Programs that are organized around the data they manage tend to be simpler, smaller, and significantly easier understand. When we do a really good job of capturing the data in our domain, the rest of the system tends to fall into place in a way which can feel like it’s writing itself. So, this book is about data. What it is, how to think about it, how to understand its semantics, how to model it, how to represent it in our code, and, somewhat surprisingly, how to listen to its feedback. The act of trying to capture what data is can often reveal how much we don’t understand about the domain we’re supposed to be modeling. To quote Bertrand Russel, “everything is vague to a degree you don’t realize until you try to make it precise.” Data-Orientation gives us the tools for making things precise. All you’ll need to follow along is a basic working knowledge of Java. As
Page
5
long as you know what a class is, and how to define an interface, and have at superficial understanding of generics (i.e. you’ve used a type like List<String> before), you’ve got everything you need to follow along. Thanks again for purchasing this book! Please share your thoughts and questions in the liveBook Discussion forum. Tell your friends and coworkers to buy a copy, too. I need the royalties to buy a yacht. —Chris Kiehl In this book welcome 1 Data Oriented Programming 2 Data, Identity, and Values 3 Data and Meaning 4 Representation is the Essence of Programming 5 Modelling Domain Behaviors 6 Implementing the Domain Model 7 Guiding the design with properties 8 Business Rules as Data 9 Refactoring towards Data 10 Data Oriented Architecture 11 Data Oriented Testing 12 Integration Testing
Page
6
1 Data Oriented Programming This chapter covers Introducing Data-Oriented programming Data as Data How representation affects our programs This book is about data. It covers what it is, how to think about it, how to represent it in our code, and all the good things that happen when we do. Programs organized around data are simpler, smaller, and easier to understand. The central idea of data-oriented programming (DoP) is modeling data “as data” in Java. That means representing it as ordinary values out in the open, liberated from the confines of private instance state. We’ll still use objects and object-orientation, but data is the building block around which we design. We’ll use it to encode business rules into the type system, make illegal states impossible to represent, draw architectural boundaries, enforce domain semantics, and create tests so powerful that they can prove entire features are correct. The goal of the book is to teach you how to apply DoP to complex, real world software services. We’ll touch on financial systems, rocket manufacturing, rules engines, and more. In some chapters we’ll build these systems from scratch; In others, we’ll drop into legacy ones and focus on refactoring. You’ll see where DoP works, where it doesn’t, and how it fits into the wider Java landscape. DoP changes what we focus on when building software. It lets us temporarily set aside object-oriented questions about “what does it do?” and instead focus on something more fundamental, bordering on philosophical, the question of “what is it?” If we learn to answer that, we unlock a powerful new way of programming.
Page
7
1.1 Objects in a Data Oriented World Since we’re all Java programmers, it’s worth clearing this up before we go any further: Data-orientation doesn’t mean giving up objects! Data orientation is not some new paradigm here to replace all the others and shame you for ever having used them. Objects are invaluable tools. I’ve tried programming without them in other languages and I always come crawling back. The only mental shift is viewing objects not as “The One True Way” to build software, but instead just another tool at our disposal. Objects are great at some things and less great at others. Such is the nature of all tools. A screwdriver can act like a chisel in an emergency, but it’s better at being a screwdriver. So, the main change is where and how much we use objects. We let objects enforce boundaries, and encapsulate state, and all the other things they’re good at, but for everything else we’ll use data. Data Orientation frees data from the shackles of private instance state. It doesn’t try to encapsulate it or hide it behind behaviors. It models it out in the open as pure information. Listing 1.1 A class modeling information class Point { #A private final double x; #A private final double y; #A // constructor, getters, equals, etc... } Classes like Listing 1.1 can feel funny at first. Some would even call it an anti-pattern due to being “anemic.” The atomic unit of object orientation is data and behavior together, coordinating as one. Further, good object- oriented design encourages us to abstract “above” the data as much as possible, and focus on the interfaces through which our objects interact. The object-oriented design process is largely the act of figuring out where you draw borders, and who calls which interface, and which piece is responsible
Page
8
for what (plus making sure each object has enough to do, but also not too much). For all the guff Java gets for being the “kingdom of nouns,” we actually spend the bulk of our effort stressing about how the verbs attached to those nouns interact with one another. Figure 1.1 Objects and how they communicate is our focus during object-oriented design The process looks much different when we focus on “data as data” during our
Page
9
design phase. The representation of that data becomes the primary point of interest. It also becomes the main point where we can improve a design. We get to take a brief step back from the complexities of how objects interact, and instead focus on the much simpler question of “what is this stuff?”. We’ll spend a lot of our time exploring what the information in a given domain is and the fundamental semantics which govern it. So where do we put the objects? My sell to you is that if we do a really good job of representing our data, the objects we need tend to naturally emerge in the right spot, often in a way that feels almost inevitable. They fall naturally into place because their role, vending a piece of data, gets figured out during our modeling phase. All that’s left is gluing everything together atop our foundation of data. It will definitely take some getting used to. There will be some challenges ahead of us. However, if we can quiet that little voice of discomfort and skepticism long enough to get over the initial hump, there’s an exciting style of programming to explore. Focusing on modeling data as data puts a very different kind of design pressure on us from the one we usually feel under object orientation. When data is out there on its own, it suddenly needs to do something that it never needs to do when encapsulated behind an object: it has to describe itself. We're forced to consider what the data really means and, most importantly, how it should be represented. It is our data’s "representation", to quote Fred Brooks, that "is the essence of programming." When we get it right, everything else falls into place. 1.2 The soul of data-oriented programming in a single line To explore the outsized effect that data’s representation has on our programs, we only need a single line. We'll give ourselves a sole piece of data. An identifier of some kind. It'll be out there on its own. Unadorned and without any kind of containing object. Listing 1.2 One line. A vague identifier of some kind String id; #A
Page
10
So, the question is, without a class to contextualize it, just what is this “id” thing? What does the current representation of this data communicate to us as readers of the code? I’d argue “not much.” When we encounter this code, all we know is that it involves a variable called "id," and that it's a String, and that’s it. The representation doesn’t tell us anything about what the id is supposed to be. A string can represent just about anything. So, we’re forced to figure out, of that “just about anything” that it could be, what should it be? What is a meaningful identifier in our domain? More often than not, this domain information doesn’t live in the code itself. Instead, we have to chase it down elsewhere -- either a more tenured coworker, external docs or wikis, or (on many projects I’ve worked on) watching production traffic to see what flows across the wires. And this is the core of the problem. The current representation doesn't communicate anything to us about what this piece of data is supposed to be. Even if we give a better name than id, or sprinkled in descriptive comments, the representation still betrays the semantics. The current representation is a little speed bump in the code that slows down each person's understanding of what's going on. It's small and trivial in this example, but small ambiguities grow into big ones. This is an ambiguity that people will have to spend time resolving. If they do it right, the only cost is time. If they do it wrong, the cost is usually a bug. So, here's the data-oriented view of this: when we're designing our data type, which, in this case, represents an identifier of some kind, we'd look at that id and ask: "Ok. What does it really mean to be an identifier in our domain?" It's probably something with more specific semantics than String. Another way of thinking about it would be to imagine if we were to drop someone fresh into this part of the code. What is it that we'd want them to know about this id? How can we communicate what we know directly in the code? As for our identifier, it could be anything, but let’s say that they’re UUIDs. Now we know what it is, we can ask if the code captures the semantics of
Page
11
"being" a UUID. Our String, while it can technically hold a UUID in string form (and is extremely common to do so), can also hold things that aren't UUIDs. Listing 1.3 One line. Many things that aren’t valid IDs String id = "not-a-valid-uuid"; #A id = "Hello World!"; #A id = "2024-05-04"; #A id = "172.16.24.105"; #A id = “1010011001011011”; #A Our one-line example, while about as trivial as it gets, embodies the problems that an imprecise representation can cause in our programs. There's a mismatch between what we "know" something means ("This id is a UUID") and what the code says it is ("literally any String is A-OK"). There's an infinite number of things that aren’t UUIDs, and our representation allows every one of them to be incorrectly assigned to our id! “An infinite number of wrong ways to do something” is a heavy burden to bear. The standard approach to this problem is usually trying to wrangle those illegal states under control with defensive programming. A precondition here, an extra if check there. Regardless of how we do it, we have to do it somewhere, because the representation allows those wrong states to exist. And mounting those defenses means writing more code. And then tests of that code. And then more code after that. All to work around the fact that the representation we've picked for something that's supposed to only be a UUID allows things that aren't UUIDs. That brings us back to the data-oriented view. All these problems stem directly from how our data is represented in the code. A String is not a good way to represent a UUID. It allows too many things that aren’t UUIDs. We know what our data is supposed to be, so we need to bring the representation of the data in closer alignment with it. In the case of our UUID, this is super simple, because Java has a ready-made type we can use. We can swap out the ambiguous String for the concrete UUID.
Page
12
Listing 1.4 Changing how we represent that one line String id UUID id #A Which is an obvious change, right? (“obvious” things that aren’t common currency are a big theme of this book). However, what's interesting about changing this one single line, as "obvious" as it might be, is that it fundamentally changes what the code communicates to us as readers. We’ve made the code describe itself. There’s no ambiguity. We don’t have to chase down coworkers or dig through databases. What it means to be an id in our domain is expressed directly in the code. This subtle shift in representation, despite being a single line, has a massive impact on our program. We've moved the semantics out of our heads and into the code where it belongs. As a result, all of the illegal “not-UUID” states disappeared! Actually, something even more profound happened: our program can now only construct correct states. There is literally no way to create anything that isn't a valid UUID. As such, there are no bad states to defend against, so there’s no need for additional code. And no need for tests of that additional code. It’s hard to quantify because it’s invisible, but picking a better representation made our entire code base smaller and simpler because of all the other code we now don’t have to write. There’s one more thing that’s really interesting about this little “obvious” example: how infrequently we collectively take this incremental step towards making our code express itself. At the time of this writing, if you search for "String id" on Github, you'll find 4.5 million instances of this exact ambiguity in Java alone. Some of those might very well actually mean “an ID is literally any conceivable String,” however, I suspect that most of those represent a gap in the developer’s modelling -- an instance where a just a little bit of unnecessary friction gets placed on the programmer's ability to understand the code. Data Orientation is largely just taking that tiny incremental extra step towards the "obvious" representation. 1.3 Show me your data, and the rest will be obvious Show me your flowcharts and conceal your tables, and I shall continue
Page
13
to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious Fred Brooks One of the central ambiguities in Java code is what the state inside of our objects means. We assume our readers will just “get it” while reading our code in the same way we do while writing it. However, so often, there’s a mismatch between what we think we’ve expressed in the code, versus what we’ve actually expressed in the code. This presents as explanations like "well, if this one field is set, then it means this, but if this other field is set then it means..." (often times this keeps going "if they're both not set then it means... " (and going “but when we set this one…”)). The state inside our objects is being used to model… something, but that “something” is only alluded to and hinted at. It’s never expressed directly. It’s fine when we’re the ones writing the code and have the implicit meaning loaded in our head. However, it becomes impenetrable when we’re the ones piecing together why that one method sets that one field some of the time. Here's an example. Listing 1.5 is a small slice of a larger set of classes that deals with running scheduled tasks asynchronously. Our job is to figure out what the reschedule method does. Listing 1.5 if attempts is not set then it means..." class ScheduledTask { private Instant scheduledAt; #A private int attempts; #A public void reschedule() { #B if (this.someSuperComplexCondition()) { this.scheduledAt = now(); this.attempts += 1; } else if (this.someOtherComplexCondition()) { this.scheduledAt = now().plusSeconds(DELAY); this.attempts = 0; } else { this.scheduledAt = null;
Page
14
this.attempts = 0; } } } The code itself, we could all probably agree, is trivial. The statements are all easy to follow, and there aren't too many of them. But easy to read doesn’t always mean easy to understand! If you try to reason about what this code does, you'll find that something important is left unsaid: what does all of this do? I don't mean what fields does it set, that's obvious. What’s not obvious is what it means when those fields are set. The challenge with understanding code like this is that you have to take off your engineering hat and swap it out for something more like a psychologist’s. We have to peer inside the dark recesses of the original developer's mind to piece together what they were thinking. Those fields mean something when they take on (or don't take on!) certain values. The author of this code had a model in their head, but they didn’t express it in the code. Of course, it's not totally grim. Dealing with this type of ambiguity might as well be what programming is most of the time. It’s what we do every day. If you stare at this code long enough you could start to make some decent guesses as to what the various branches mean. For instance, in the first branch, where someComplexCondition is true, attempts gets incremented. However, that same attempts variable is reset in all other branches. Pair those together and you have interesting clues that there’s some kind of lifecycle going on here. It’s unstated in the code, but we can see it in between the lines if we squint. Listing 1.6 Looking closer at the reschedule method public void reschedule() { #A if (this.someSuperComplexCondition()) { this.scheduledAt = now() this.attempts += 1; #A } else if (this.someOtherComplexCondition()) { this.scheduledAt = now().plusSeconds(DELAY); this.attempts = 0; #B } else {
Page
15
this.scheduledAt = null; this.attempts = 0; #B } } When it comes to understanding what these variable assignments mean, the absolute best we can do within the confines of this method is make a guess. There's just not enough information to build up a working theory of why these values are set or what it means when they are. We’ve reached an informational dead end because the code isn’t expressive enough to communicate. So, we have to leave the method and go foraging around in the wider world. The code doesn’t tell us what these variable assignments mean, so we have to look at where and how they’re used. It's only by studying everything that we can inductively start to build up our own mental model that explains what these individual variable assignments denote. If you're lucky, you'll find usages elsewhere that give context as to what a particular assignment means. Listing 1.7 gives some meaning to one of the possible states. Listing 1.7 finding clues throughout the code class Scheduler { private void purgeCancelledTasks() { this.tasks().removeIf((task) -> task.scheduledAt() == null #A } } The process is to just keep going like this, over and over, poking around the codebase, reading all of the implementations, until we’ve found a large enough set of examples to construct our own model of what assignments in the code mean. Table 1.1 the meaning we’ve worked out for various states assigned in the reschedule method
Page
16
When these are set Then it seems to mean… attempts+=1; scheduledAt = small delta Go for an immediate retry attempts=0; scheduledAt=large delta; Give up for now and try again later attempts=0; scheduledAt=null Give up on this job entirely This process always represents hard won knowledge. It could be a few minutes, hours, or even days depending on how complex the code is. And you can never really be sure that your model of the world lines up with the one originally envisioned. There could be another usage lurking in the code that invalidates our assumptions. It's similar to the old observation about the limits of software testing: you can only prove the existence of bugs, never their absence. Understanding an existing codebase written in this style presents us with the same limitations. It's a lossy process. Without a direct line to the original developer, we can only reason inductively about the examples we’ve found. We can’t be sure, because the code doesn’t tell us. This lack of explicit and clear modeling presents a very real problem. Beyond being hard to understand, I’d argue that it’s one of the most common sources of bugs in our programs. To use some lingo from the world of databases, the code in our example lacks “semantic integrity.” Those variables mean something when they’re set, and the interpretation of that meaning is critical to the correct functioning of the software, but that meaning isn’t enforced or encoded anywhere. It’s left entirely implicit, which means each and every part of the code has to reinterpret that implicit meaning over and over again. This dooms us to semantic drift as time goes on. All it takes is for one method to misinterpret what a particular state means for the whole thing to collapse in on itself. Which all brings us back to the data-oriented view. There’s some data in there that’s floating just below the surface. It’s the ideas behind what these individual variable assignments mean. It’s not currently expressed, but we can bring it to light if we focus on it. This is where representing data “as data” comes in. Focusing on just the data by itself lets us take a step back from the details of the code and instead just think about, ok, what is it that I'm really talking about? What does it mean
Page
17
when I set those fields? I've got some mental model floating around in my head, and it has a semantics that I'm implicitly honoring. How do I put what I know into the code? So, what are we really talking about in this case? When a ScheduledTask fails, our system has to make a decision about what to do next. What we’ve been piecing together as we explored this code is that it’s not just any arbitrary decision, it selects from a fixed set of possible decisions. That’s what’s those variable assignments mean. It’s just hidden by the current modeling. Figure 1.2 Being explicit about what a task can transition to after failing
Page
18
So, let’s start there. Let’s model these decisions that our system can make as individual pieces of data. Modern Java gives us lots of options for how to do this, but for now we’ll just use the humble class. But our intention with this class is not to create an object. We won’t be defining any behaviors. These classes will represent data. They’ll be defined entirely by their name and their attributes. Nothing else. (Technically, these classes won’t quite be “data” as we define it later in the book, but we’ll hand wave that away for now. Exploring how we get objects to act like data is the topic of the next chapter). Figure 1.3 Representing each decision as a piece of standalone data
Page
19
Look how descriptive those data types in Figure 1.4 are. They don’t do anything – they’re just data about a behavior that the system is allowed to do elsewhere. But despite just sitting there being data, they tell us so much about important ideas within our domain. We can understand what happens after a task fails just by looking at our data. We don’t have to guess or interpret vague state assignments. It’s obvious and explicit! Let’s plug it into the original code. Listing 1.8 Refactoring to make the semantics explicit
Page
20
class ScheduledTask { private Instant scheduledAt; #A private int attempts; #A private RetryDecision status; #A public void reschedule() { if (this.someSuperComplexCondition()) { this.status = new RetryImmediately( #B now(), this.attemptsSoFar(status) + 1 } else if (this.someOtherComplexCondition()) { this.status = new ReattemptLater( #B now().plusSeconds(DELAY) } else { this.status = new Abandoned(); #B } } Don't worry about how we tie all these data types together for now (we'll get to that in later chapters). Instead, the thing we should pay attention to is the fact that by splitting out the data on its own and forcing it to carry its own weight, what the code as a whole communicates has been completely transformed. This change is very similar to our one-line change in section 1.2. It’s ultimately pretty small and obvious, maybe even underwhelming at first. However, its effects ripple throughout our codebase. You can now read this one method and understand a lot about the behavior of the system. We can understand something fundamental about the core ideas in our domain just from looking at the data types. No more guess work. No more hunting for clues. The code tells you exactly what it is. Compare that to the variable assignments we had before. Table 1.2 Comparing the meaning conveyed with different approaches Old implicit variable assignments Explicit concrete Data attempts+=1; class RetryImmediately {