Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Brian Godsey

Rating No ratings yet

Data collected from customers, scientific measurements, IoT sensors, and so on is valuable only if you understand it. Data scientists revel in the interesting and rewarding challenge of observing, exploring, analyzing, and interpreting this data. Getting started with data science means more than mastering analytic tools and techniques, however; the real magic happens when you begin to think like a data scientist. This book will get you there. Think Like a Data Scientist teaches you a step-by-step approach to solving real-world data-centric problems. By breaking down carefully crafted examples, you’ll learn to combine analytic, programming, and business perspectives into a repeatable process for extracting real knowledge from data. As you read, you'll discover (or remember) valuable statistical techniques and explore powerful data science software. More importantly, you’ll put this knowledge together using a structured process for data science. When you've finished, you'll have a strong foundation for a lifetime of data science learning and practice. What’s Inside • The data science process, step-by-step • How to anticipate problems • Dealing with uncertainty • Best practices in software and scientific thinking Readers need beginner programming skills and knowledge of basic statistics. Brian Godsey has worked in software, academia, finance, and defense and has launched several data-centric start-ups.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, process-first guide for aspiring data scientists who want to move beyond tools and techniques, teaching a repeatable mindset for turning messy real-world data into actionable answers. Ideal for beginners with basic programming and statistics who want a structured approach to the entire data science lifecycle. 【Book Arc】 - **Opening (~0%–9%)**: Establishes the core philosophy that data science sits between statistics and software, arguing that *thinking* like a data scientist matters more than mastering any specific tool. Introduces the book's goal: a repeatable process for real-world problems, not just theory. - **Early (~9%–28%)**: Focuses on the foundational first step—setting goals by asking good questions. Emphasizes that analysis demands a question, that customers are rarely data scientists, and that domain knowledge from nontechnical colleagues is essential. Includes practical guidelines for coding patterns and avoiding common pitfalls like pride or shyness. - **Early–Middle (~28%–38%)**: Delves into planning and anticipation. Uses a vivid example of a colleague misapplying mean reversion to show how statistical misconceptions can derail projects. Stresses thinking through software choices, data formats, and potential obstacles (missing data, too much data) before starting. - **Middle (~38%–47%)**: Explores the "virtual wilderness" of data—where data lives, its forms (flat files, XML, JSON, databases, APIs), and how to interact with it. Frames data science as applying the scientific method (ask, hypothesize, predict, test, conclude) to digital data, with a historical look at how the internet and IoT created today's data-rich environment. - **Late (~47%–end)**: Moves into building a product with software, covering data wrangling, assessment, and translating statistics into code. The excerpts show a progression from conceptual planning to practical implementation, with emphasis on using built-in methods versus writing custom ones. 【Key Takeaways】 - **Thinking beats tools** (Opening): The author's central claim is that a data scientist's thought process—asking the right questions, framing problems correctly—matters more than any specific software or statistical technique. This reframes learning from "mastering tools" to "mastering a mindset." - **Every analysis demands a question** (Early): A recurring theme is that "Can you analyze this for me?" is never that simple. You must define a clear question and goals before touching data, or the work risks being wasted. This is the foundation of the entire process. - **Customers are not data scientists** (Early): Expect customer expectations to be unclear or inappropriate; treat goal-setting as a joint exercise, almost like conflict resolution. The data scientist and customer have different perspectives, and finding agreement is the first project milestone. - **Domain knowledge is non-negotiable** (Early): Pride or shyness prevents you from learning from experts in the field (e.g., a genetics professor). Their input on goals and caveats (like excluding certain microRNAs) is critical to avoid adding noise to your analysis. - **Beware statistical misconceptions** (Middle): The mean-reversion example (a colleague betting on a model's success rate "returning" to average) illustrates a common error. Not all systems revert to the mean; a fair coin has no memory. Always question whether your assumptions match the data-generating process. - **Plan software choices deliberately** (Middle): Think through data format, transformations, data volume, and loading methods *before* choosing a tool. Your favorite tool may not fit the problem; deliberate planning prevents costly mistakes. - **Data is a wilderness to explore** (Middle): Treat data science as applying the scientific method—ask, hypothesize, predict, test, conclude—to digital data. The internet and IoT have created a vast, messy "wilderness" of data that requires this structured exploration. 【Reading Tips】 - **Deep-read Chapters 1–2** (Early): These set the philosophical and practical foundation. Pay special attention to the goal-setting framework and the "customer is not a data scientist" discussion—they'll inform every later step. - **Skim the historical data-growth section** (Middle ~44%): The story of how the internet and IoT created today's data landscape is interesting but not critical for practice. Focus instead on the scientific method application and data formats discussion that follows. - **Use the mean-reversion example as a cautionary tale** (Middle ~34%): This is a memorable, concrete illustration of a statistical trap. Use it to internalize the habit of questioning assumptions about data-generating processes. - **Treat the software chapters as reference, not scripture** (Late): The author deliberately keeps software coverage abstract to stay relevant. Skim for concepts (e.g., built-in vs. custom methods) rather than memorizing specific tools, which will change over time. - **Take away the process, not the tools**: The book's value is the repeatable step-by-step approach. As you read, build your own checklist from the guidelines (e.g., anticipate obstacles, think through software choices) to apply to future projects. 【Coverage Limits】 The excerpts primarily cover the book's early-to-middle sections (philosophy, goal-setting, data exploration). Detailed content on later stages—data wrangling specifics, assessment techniques, and software implementation—is only partially represented, so this guide's depth on those topics is limited.
Page 7
ng and prodding 84 PART 2 BUILDING A PRODUCT WITH SOFTWARE AND STATISTICS ........................................................105 6 ■ Developing a plan 1...
View in text
Excerpt 2
n these fields. But I despise jargon and presumed knowledge more than most, and so I’ll try hard to include accessible conceptual explanations of statistical...
View in text
Excerpt 3
ding agreement between two personalities, two perspectives, that if they aren’t conflicting are at the very least disparate. Although there is not, strictly...
View in text
Excerpt 4
similar-sized collections of email addresses were no longer 44 CHAPTER 3 Data all around us: the virtual wilderness This idea of data as a wilderness is one...
View in text
Excerpt 5
ely mistaken in creating the data format in the first place. When in doubt, send a few emails and try to find someone who can help you. 3.2.10 Deciding which...
View in text
Excerpt 6
ata is preceded by a <PRE> tag. Regardless of what this tag means, it’s worth checking to see if it appears often on the page or only right before 78 CHAPTER...
View in text
Excerpt 7
n. Then, dissect your description, looking for assumptions. For example, I might describe my original project involving the Enron data like this: “My data se...
View in text
Excerpt 8
analysis tech- niques to try to detect suspicious behavior. This was an open-ended project. We knew that some bad, criminal things hap- pened at Enron, but w...
View in text
Tags
AI categories
data scienceProgrammingTechnology
ISBN: 1633430278
Publish Year: 2017
Language: English
Pages: 307
File Format: PDF
File Size: 5.2 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…