Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Tommaso Teofili

Rating No ratings yet

Deep Learning for Search teaches readers how to leverage neural networks, NLP, and deep learning techniques to improve search performance. Deep Learning for Search teaches readers how to improve the effectiveness of your search by implementing neural network-based techniques. By the time their finished, they'll be ready to build amazing search engines that deliver the results your users need and get better as time goes on!

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Deep Learning for Search — Reading Guide ## 【One-Line Pitch】 A practical handbook for search engineers and programmers who want to apply neural networks, NLP, and deep learning techniques to build smarter, more relevant search engines that improve over time. If you know search fundamentals but are new to deep learning, this book bridges both worlds with hands-on Lucene-based examples. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces the promise of neural search — why deep learning matters for search, from image retrieval without manual metadata to cross-language search and direct-answer responses. Sets expectations that no prior DL knowledge is required. - **Early (~10%–23%)**: Covers search fundamentals: text analysis (tokenizers, token filters, analyzers), indexing with inverted indexes, and the concept of relevance. Explains how deep neural networks work — layers, neurons, weights — and why compositional data like text and images benefits from deep architectures. - **Early (~23%–32%)**: Discusses neural network training workflows and the tension between static ML models and continuously updating search indexes. Introduces the practical challenge of retraining models as new content arrives. - **Middle (~39%–48%)**: Dives into synonym expansion as a recall-improvement technique, using Lucene's `SynonymGraphFilter` with concrete code examples. Builds the intuition that context matters — leading naturally to learning word representations from data rather than hand-crafted vocabularies. - **Middle (~48% onward)**: Transitions from manual synonym dictionaries to learned word embeddings, using the library-assistant analogy (John vs. Robbie) to explain why data-driven synonym discovery beats static linguistic knowledge. Later chapters (per table of contents) cover suggesters, ranking with word embeddings, document embeddings for recommendations, and cross-language search. ## 【Key Takeaways】 - **Deep learning removes manual metadata bottlenecks** (Early): Neural networks can abstract representations of images and text directly, eliminating the need for human-typed descriptions before indexing. This is the core motivation for neural search. - **Search fundamentals still matter** (Early): Tokenization, filtering, and inverted indexes remain the backbone of any search engine. Understanding these basics is prerequisite to applying DL techniques effectively. - **Deep networks suit compositional data** (Early): Text and images decompose hierarchically (paragraphs → sentences → words), making them ideal candidates for deep architectures that learn incremental representations. - **Static models clash with dynamic indexes** (Early): Standard ML workflows assume fixed training sets, but search engines ingest continuous streams of new content. You must plan for retraining or incremental model updates when architecting neural search. - **Synonym expansion improves recall** (Middle): Expanding query terms with synonyms at index or query time lets users find relevant documents even with imprecise wording — like searching "music is my plane" and retrieving "music is my aeroplane." Recall is defined as retrieved-relevant divided by total-relevant documents. - **Context is the key to word meaning** (Middle): Words that appear in similar contexts (e.g., "live" and "visit" near places) carry related meanings. This intuition underpins learning word representations from data rather than relying on static vocabularies like WordNet. - **Data-driven synonym discovery beats linguistic expertise** (Middle): A domain expert who reads the actual corpus (Robbie) outperforms a grammarian with formal knowledge (John) when it comes to recognizing domain-specific synonyms like "AI" = "artificial intelligence." This motivates embedding-based approaches. ## 【Reading Tips】 - **Skim the opening chapters** (~0–10%) if you already know search basics; they're introductory and set context but don't contain advanced techniques. - **Deep-read the Lucene code examples** (Middle, ~39–48%): The `SynonymGraphFilter` configuration and analyzer setup are concrete and reusable. Follow along with the code listings to understand index-time vs. search-time expansion. - **Pay attention to the library-assistant analogy** (~48%): It's the conceptual bridge from manual synonyms to learned embeddings — the pivotal moment in the book's argument. - **Watch for the retraining problem** (Early, ~32%): This is a recurring practical concern; note how the author returns to it as later chapters introduce more complex models. - **Use the table of contents as a roadmap**: Later parts cover suggesters, ranking, document embeddings, and cross-language search — skim ahead if your interest is specific to one of these areas. ## 【Coverage Limits】 This guide synthesizes excerpts from roughly the first half of the book (through ~48%). Later chapters on suggesters, ranking with word embeddings, document embeddings, and cross-language search are listed in the table of contents but not covered in detail here. ##
Excerpt 1
age search 199 ■ Querying in multiple languages on top of Lucene 200 acknowledgments First and foremost I would like to thank my lovely wife Michela for enco...
View in text
Excerpt 2
rch engine serving both text and images seamlessly to users from all over the world, which, instead of returning search results, returns the single piece of...
View in text
Excerpt 3
“When and Why Are Deep Networks Better Than Shal- low Ones?” Proceedings of the AAAI-17: Thirty-First AAAI Conference on Artificial Intelligence (Center for...
View in text
Excerpt 4
ng that, let’s look at how song lyrics are structured. Each song has an author, a title, a publication year, lyrics text, and so on. As I said earlier, it’s...
View in text
Excerpt 5
ccess the original data as it existed before it was indexed. In the case of indexing the top 100 songs of the year to build a search engine of song lyrics, y...
View in text
Excerpt 6
artificial intelligence” and “books about machine learning.” (This is a simple example: the two sequences are exactly the same length.) One of the first thin...
View in text
Excerpt 7
t h a l r t/7 h e Figure 4.6 A finite v/4 state transducer analyzer to the dictionary entries are compiled into a big FST. At query time, traversing the FST...
View in text
Excerpt 8
ponds to a certain one-hot-encoded vector (and vice versa): public class CharLSTMNeuralLookup extends Lookup { private CharacterIterator characterIterator; p...
View in text
Tags
AI categories
Programmingsearchdeep learning
ISBN: 1617294799
Publish Year: 2019
Language: English
Pages: 355
File Format: PDF
File Size: 7.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…