Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Ken Youens-Clark

Rating No ratings yet

Life scientists today urgently need training in bioinformatics skills. Too many bioinformatics programs are poorly written and barely maintained--usually by students and researchers who've never learned basic programming skills. This practical guide shows postdoc bioinformatics professionals and students how to exploit the best parts of Python to solve problems in biology while creating documented, tested, reproducible software. Ken Youens-Clark, author of Tiny Python Projects (Manning), demonstrates not only how to write effective Python code but also how to use tests to write and refactor scientific programs. You'll learn the latest Python features and tools--including linters, formatters, type checkers, and tests--to create documented and tested programs. You'll also tackle 14 challenges at Rosalind, a problem-solving platform for learning bioinformatics and program

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide for life-science researchers and students who can already write a little Python but need to produce documented, tested, reproducible bioinformatics software. It teaches professional coding habits—tests, linters, type checkers, and clean structure—by working through real Rosalind-style biology challenges. 【Book Arc】 - **Opening (~0%–10%)**: Sets up the book's premise—bioinformatics code is often fragile and unmaintained—and explains the two-part structure: 14 Rosalind challenges followed by more complex programs, each with a test suite. - **Early (~10%–32%)**: Establishes the core workflow through small challenges like counting nucleotides and transcribing DNA, introducing named tuples, dictionaries, `dict.get()`, and the test-driven cycle of writing specs as executable tests. - **Middle (~32%–48%)**: Moves into file processing and algorithmic thinking—reading/writing directories of sequences, reverse complements, and benchmarking solutions like Fibonacci—while layering in regular expressions, list comprehensions, and Biopython. - **Late (~48%–70%)**: Tackles harder problems such as protein motif search with regex, mRNA inference from protein, and restriction-site analysis, emphasizing code sharing, testing, and alternative solution strategies. - **Ending (~70%–100%)**: The second part presents more complex programs demonstrating additional patterns and concepts important in bioinformatics; the excerpts do not cover these chapters in detail. 【Key Takeaways】 - **Tests are specifications made executable** (Early): The author frames pytest suites as the definition of "done," writing tests before solutions and running them after every change—this is the book's central discipline. - **Write the happy path, avoid exception clutter** (Early): Following Joe Armstrong's Erlang philosophy, the book deliberately avoids try/catch complexity in short research programs, letting errors crash rather than obscuring logic. - **Multiple solutions beat one "obvious" way** (Early): Each chapter uses a theme-and-variations approach, showing several implementations to explore different Python syntax and data structures—valuable for learning, even if it contradicts the Zen of Python. - **Tooling enforces quality** (Middle): Linters (pylint, flake8), type checkers (mypy), and formatters are integrated into the workflow, with a `make test` shortcut aiming for a completely clean suite. - **Data structures shape solutions** (Early): Named tuples make records ergonomic, dictionaries serve as lookup tables and decision trees, and lists double as stacks—choosing well simplifies the code. - **Real bioinformatics means file plumbing** (Middle): Processing directories of sequencing files, creating output directories, and handling one-or-many inputs is treated as a core pattern, not an afterthought. - **Benchmarking is part of algorithm work** (Middle): The Fibonacci chapter explicitly compares implementations, teaching readers to think about performance, not just correctness. - **Biopython is a legitimate endpoint** (Middle): After writing algorithms by hand for learning, the book shows library solutions like `Bio.Seq`, signaling when to stop reinventing. 【Reading Tips】 - **Deep-read the test files**: The tests are where the author's real teaching happens—study `tests/*_test.py` to understand how requirements become assertions. - **Attempt each challenge before reading solutions**: The author explicitly says you gain the most by writing working programs first, then comparing your approach to the variations shown. - **Skim the setup/installation prose, slow down on solutions**: Early chapters spend time on environment and CLI conventions; the conceptual payoff is in the solution discussions. - **Treat the second part as a pattern catalog**: Since excerpts don't detail it, approach those chapters looking for reusable design patterns rather than specific biology results. - **Run the tools, don't just read about them**: Actually invoke pytest, pylint, and mypy on your own code to internalize the quality cycle. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book; the second part's complex programs and later chapters are only referenced, so specific techniques there are not summarized.
Page 9
233 Writing a Regular Expression to Find the Motif 235 Solutions 237 Solution 1: Using a Regular Expression 237 Solution 2: Writing a Manual Solution 239 Goi...
View in text
Excerpt 2
have the same features, such as generating usage statements and properly validating arguments. Create your dna.py program in the 01_dna directory, as this co...
View in text
Excerpt 3
he None value: >>> type(counts.get('N')) <class 'NoneType'> I can use the == operator to see if the return value is None: >>> counts.get('N') == None True bu...
View in text
Excerpt 4
the reverse complement, so I’ll give it a string: $ ./revc.py AAAACCCGGT ACCGGGTTTT As the help indicates, the program will also accept a file as input. The...
View in text
Excerpt 5
ion starts ============================ tests/cgc_test.py::test_exists PASSED [ 20%] tests/cgc_test.py::test_usage PASSED [ 40%] 114 | Chapter 5: Computing G...
View in text
Excerpt 6
ction, so I’ll use list() to coerce the values in the REPL: Solutions | 145 >>> seq1, seq2, = 'GAGCCTACTAACGGGAT', 'CATCGTAATGACGGCCT' >>> sum(starmap(operat...
View in text
Excerpt 7
/solution3_functional.py,./solution4_kmers_functional.py,\ ./solution4_kmers_imperative.py,./solution5_re.py \ '{prg} GATATATGCATATACTT ATAT' --prepare 'rm -...
View in text
Excerpt 8
() and max() functions that will return the minimum or max‐ imum value from a list: >>> min(map(len, seqs)) 5 >>> max(map(len, seqs)) 7 So the shortest seque...
View in text
Tags
AI categories
PythonProgrammingData
ISBN: 1098100883
Publisher: O'Reilly Media
Publish Year: 2021
Language: English
Pages: 456
File Format: PDF
File Size: 10.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…