Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Reuven M. Lerner

Rating No ratings yet

Practice makes perfect pandas! Work out your pandas skills against dozens of real-world challenges, each carefully designed to build an intuitive knowledge of essential pandas tasks. In Pandas Workout you’ll learn how to: • Clean your data for accurate analysis • Work with rows and columns for retrieving and assigning data • Handle indexes, including hierarchical indexes • Read and write data with a number of common formats, such as CSV and JSON • Process and manipulate textual data from within pandas • Work with dates and times in pandas • Perform aggregate calculations on selected subsets of data • Produce attractive and useful visualizations that make your data come alive Pandas Workout hones your pandas skills to a professional-level through two hundred exercises, each designed to strengthen your pandas skills. You’ll test your abilities against common pandas challenges such as importing and exporting, data cleaning, visualization, and performance optimization. Each exercise utilizes a real-world scenario based on real-world data, from tracking the parking tickets in New York City, to working out which country makes the best wines. You’ll soon find your pandas skills becoming second nature—no more trips to StackOverflow for what is now a natural part of your skillset. About the book Pandas Workout is a thoughtful collection of practice problems, challenges, and mini-projects designed to build your data analysis skills using Python and pandas. The workouts use realistic data from many sources: the New York taxi fleet, Olympic athletes, SAT scores, oil prices, and more. Each can be completed in ten minutes or less. You’ll explore pandas’ rich functionality for string and date/time handling, complex indexing, and visualization, along with practical tips for every stage of a data analysis project. About the reader For Python programmers and data analysts.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on, exercise-driven course for Python programmers and data analysts who want to build fluent, professional-level pandas skills through 200 realistic, ten-minute challenges—from cleaning messy data to visualizing results. 【Book Arc】 - **Opening (~0%–9%)**: Introduces the book’s structure—each chapter mixes short explanations with exercises, solutions, and "Beyond the exercise" extras—and establishes the core workflow: Jupyter notebooks, f-strings, and essential Series operations like `mean`, `max`, and `idxmin`. - **Early (~9%–25%)**: Builds foundational Series skills: broadcasting arithmetic, boolean mask indexing for filtering and assignment, fancy indexing, descriptive statistics, and `value_counts` for frequency analysis—all through small, concrete scenarios like test scores and passenger counts. - **Early (~25%–34%)**: Moves to DataFrames, covering four ways to construct them (lists of lists, lists of dicts), row/column selection with `.loc`, en masse assignment, and practical tasks like calculating revenue and finding outliers using IQR on taxi data. - **Middle (~34%–47%)**: Tackles importing and exporting data, with a deep dive into CSV quirks (no formal spec), dtype handling for memory efficiency, reading from URLs via `requests` and `StringIO`, and combining boolean conditions for complex queries on real NYC taxi data. - **Late (~47%–100%)**: Covers advanced topics—indexes and multi-indexes, cleaning (duplicates, missing values), grouping/joining/sorting (two chapters), strings, dates, visualization with pandas and Seaborn, performance optimization, and two large capstone projects (Python developer survey, American colleges). 【Key Takeaways】 - **Series are the atomic unit of pandas** (Early): Master broadcasting (e.g., `s + (80 - s.mean())`), boolean masks for filtering/assignment (`s.loc[s <= s.mean()] = 999`), and fancy indexing (`s.loc[[2,4]]`) to handle most data manipulation tasks efficiently. - **Descriptive statistics reveal data shape quickly** (Early): Use `mean`, `median`, `std`, quantiles, and `idxmax`/`idxmin` to summarize distributions; `value_counts` is a favorite for frequency tables—essential for exploratory analysis. - **DataFrame construction has four idiomatic patterns** (Early): Choose between lists of lists (positional), lists of dicts (keyed), and other methods depending on your source; `.loc` with row/column selectors enables precise retrieval and bulk assignment. - **CSV is the de facto standard but has no formal spec** (Middle): Handle quirks by specifying `usecols`, `header=None`, and explicit `dtype` dicts (e.g., `np.int8` for memory savings); remember integer columns can’t hold NaN—convert after cleaning. - **Boolean conditions combine with `&` for powerful queries** (Middle): Filter on multiple criteria (e.g., below-average distance AND above-average cost) using `loc` with parenthesized masks; `.count()` on the result gives quick tallies. - **Read data from URLs directly** (Middle): `pd.read_csv` works with web links; for non-direct URLs, fetch with `requests`, wrap in `StringIO`, then parse—enabling automated reports like Bitcoin price summaries. - **Indexes and multi-indexes are key to fluent pandas** (Late): Hierarchical indexes allow searching on parts of a hierarchy; pivot tables and multi-index manipulation are essential for advanced grouping and reshaping. 【Reading Tips】 - **Skim the reference tables** (Early): Each chapter opens with a "What you need to know" table—scan these for quick API refreshers, then dive into exercises that apply them. - **Deep-read the "Working it out" sections**: These explain the reasoning behind solutions, not just code—crucial for building intuition, especially for boolean masks and dtype handling. - **Do the "Beyond the exercise" challenges**: They’re unanswered but push you to generalize skills (e.g., using `query` instead of `loc`); solutions are downloadable if you get stuck. - **Watch for dtype pitfalls** (Middle): Pay extra attention to the sections on integer vs. float conversion and `low_memory=False`—these are common real-world gotchas. - **Use the Pandas Tutor links**: Each solution has an executable version online—run and modify code to experiment beyond the book’s examples. 【Coverage Limits】 This guide synthesizes the opening through the middle of the book (~47%); later chapters on cleaning, grouping, strings, dates, visualization, performance, and final projects are summarized from the table of contents but not detailed from excerpts.
Excerpt 1
ractical tips for every stage of a data analysis project. About the reader For Python programmers and data analysts. 200 exercises to make you a stronger dat...
View in text
Excerpt 2
= 999 The result? 0 999 1 999 2 999 3 40 4 50 dtype: int64 In this way, we replace elements less than or equal to the mean with 999. This technique is worth...
View in text
Excerpt 3
_distance'] < df['trip_distance'].quantile(0.25) - 1.5*iqr] df[df['trip_distance'] There are 1,889 > df['trip_distance'].quantile(0.75) + 1.5*iqr] high outli...
View in text
Excerpt 4
Bitcoin over the most recent year as of when you read this. (For that reason, your results will look different from mine, even if you use the same code.) Onc...
View in text
Excerpt 5
ti-index should be based on Year, Season, Sport, and Event. 2 Answer these questions: – What is the average age of winning athletes in summer games held betw...
View in text
Excerpt 6
gh the data frame df.pct_change For a given data frame, df.pct_change() http://mng.bz/4DBB indicates the percent- age difference between each cell and the co...
View in text
Excerpt 7
e many ways in which you can split, combine, and analyze it. In this chapter, we looked at some of the most common tasks—from grouping for analysis, to group...
View in text
Excerpt 8
e at a time, to the function proportion_of_city_precip. The return value is then a series in which the parallel rows from the input series have their new val...
View in text
Tags
AI categories
DataProgrammingTechnology
ISBN: 1617299723
Publish Year: 2024
Language: English
Pages: 442
File Format: PDF
File Size: 14.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…