Start with the one-hour workshop below to use the textbook at a staff meeting or professional-development session. The case, participant handout, timings and worked answers are supplied. Participants need no advance reading or subscription access.
The main audiences are information literacy and evidence synthesis librarians, with systems/discovery librarians as a secondary audience. The routes below connect each group’s work to suitable reading and an outcome they can demonstrate.
You may adapt these notes and the handout with credit under CC BY 4.0. See reuse guidance for the book’s third-party figures.
One hour: why did this search miss a paper?
For: academic librarians who teach searching or support research consultations. Familiarity with ordinary database searching is enough. Outcome: participants can distinguish a missing citation from a missing candidate, choose a check that would narrow the diagnosis, and explain a sensible next action to a researcher.
This is a short introduction to Application Exercise III. The full exercise adds live searching, relevance judgements and measurement; those need time beyond this hour. The timings below are a suggested first-run plan and have not yet been validated in an external teaching pilot.
Prepare the session
- Facilitator preparation: allow about 20–30 minutes to read the case and worked answers, then the book’s anatomy of a cited answer, diagnostic questions and opening of Chapter 15. Allow longer if these distinctions are new to you.
- Materials: share or print the two-page participant handout. It contains the case, response spaces and a transfer check, with no answer key. Participants can write on paper or copy the prompts into a document. This page supplies your facilitation notes; a slide deck is unnecessary.
- Access: use pairs or small groups, with individual writing before discussion. Read the case aloud as well as providing it in text. A printed handout is sufficient; the session does not depend on a live product, login or network connection.
The prepared case
Teaching scenario: this uses the book’s open-access citation-advantage question. The observations and trace below are stipulated for practice; they are not findings from a new test of a named product.
A researcher asks an academic search assistant, “Is there an open access citation advantage?” Its answer cites five papers. The researcher identifies another relevant paper that is not cited. An exact-title search in the service’s ordinary discovery mode finds that paper.
The service’s documentation describes this assistant mode: it transforms the question into a lexical query, retrieves up to 30 candidates, reranks them, and supplies up to five to the answer writer. Initially, the researcher can see only the answer and its citations.
Ask participants: what is established, what remains possible, and which additional evidence would best distinguish those possibilities? Then have them propose one next action. Finding the paper in another mode does not establish that this assistant searched the same index, fields or route.
Run of session
| Minutes | Facilitator action | Participant output |
|---|---|---|
| 0–5 | Read the case. Ask everyone to write their first explanation before discussing it. | An initial diagnosis and next action to revisit at the end. |
| 5–12 | Explain three distinctions: entering a candidate set; ranking and passing a smaller set onward; choosing what to cite in an answer. A paper can disappear at each boundary. | A simple pipeline annotated with possible exclusion points. |
| 12–25 | Pairs separate observations from hypotheses and choose one discriminating check. Ask what each proposed check could actually establish. | Two plausible explanations, with the evidence needed to distinguish them. |
| 25–35 | Reveal the trace in the facilitator answer below. Ask participants to revise their diagnosis and select an intervention. | A diagnosis tied to a specific boundary, with a testable next action. |
| 35–47 | Ask for a short explanation suitable for a researcher. Partners check that it preserves the limits of the evidence. | Three to five sentences: what happened, what to try, and what that would leave unknown. |
| 47–55 | Use the new case on page two of the handout, individually, before discussing the answer. | A diagnosis for a different exclusion point. |
| 55–60 | Return to the initial explanation. Ask each participant to identify one revision and one place they could use the method at work. | A corrected explanation and a concrete next use. |
Facilitator answers and the trace to reveal at minute 25
Before the reveal: the missing citation establishes absence from the answer’s cited set. It does not locate the failure. The paper may be absent from the assistant’s candidates, excluded by reranking and its cut-off, or supplied to the writer and left uncited. A trace of the actual run, including the candidates and writer’s input, distinguishes these. Documentation describes the intended pipeline; it is not a trace of this particular search.
Read this new evidence aloud: “A saved trace of the original assistant run is now available. The paper was among the 30 first-stage candidates, ranked eighteenth. After reranking it was eighth. Only the first five were supplied to the answer writer.”
Diagnosis: the paper entered the first-stage candidate set but did not cross the later five-paper boundary. Generation could not cite it from its supplied evidence. The trace locates the exclusion; it does not explain why the reranker placed the paper eighth.
Reasonable next actions: inspect the paper directly for the immediate research need; try a targeted reformulation or another retrieval route, preserving the original run for comparison; or test a larger downstream candidate budget if that control exists. Explain what the action tests. A new query can change the candidate set and ranking, but recovery is not guaranteed. Asking the writer to reconsider the same five papers cannot add the missing one to its supplied evidence.
Example explanation to the researcher: “The saved trace shows that the search found this paper, but it was eighth after reranking and only five papers reached the answer writer. We can inspect it directly and try a more targeted search, saving both runs to see what changes. Recovering this paper would help with your question, but would not establish that we had found every relevant study.”
Transfer-check answer: in the second case, the paper reached the writer’s input. Its absence from the answer therefore points to generation or citation selection, rather than exclusion by the retrieval stages in that run. Compare the relevant source passage with the answer and inspect any available citation-selection record. Inclusion in the input does not prove that the writer used it faithfully; one recovered paper does not establish complete retrieval.
Look for: a revised diagnosis after new evidence, an intervention suited to the boundary, and an explicit limit on the conclusion. If participants keep treating a missing citation as proof of failed retrieval, return to the writer’s input before adding terminology.
After the session: invite participants to apply the same observation–hypothesis–check–action record to one real enquiry. Use the full exercise when the aim is to compare retrieval quality, and the short teaching record to improve the next delivery.
What the book covers
The book is about the retrieval half of a search system: how text becomes searchable, how records become candidates, how they are ranked, how the pieces are assembled into a pipeline, who decides what to search next, how retrieval fails, how you would measure it, and what a library must record and require. It assumes no prior information-retrieval knowledge and no mathematics — formulas and worked examples are accompanied by explanations in words. No calculation is required to follow the main argument.
Four things are out of scope by design, and it is worth saying so to a class that will expect them:
- How language models generate text. The book stops at the point where a model writes prose for a reader.
- Detecting fabricated citations. That is a generation problem; the book is about what was never retrieved.
- Prompt engineering. Not covered anywhere.
- General agent architectures and planning. One narrow slice appears — how an agent selects, repeats and stops retrieval actions — because that slice is retrieval design.
The running examples deliberately use recent internet slang — delulu, rizz, cooked, touch grass. This is not decoration: words a system has never seen, or has seen only in another sense, are where retrieval breaks in instructive ways. Some students will find it either funny or grating, and it is worth naming the reason in advance.
Choose the emphasis for the cohort
Choose the professional task first, then assign the supporting reading. These routes use the same book; a short session needs only the passages relevant to its outcome.
| Audience | Task to practise | Evidence of learning |
|---|---|---|
| Information literacy | Explain an unexpected result and choose the next search action. | A short explanation a student could use, with observation separated from hypothesis. |
| Evidence synthesis | Distinguish finding candidates from prioritising an existing screening pool. | A search or screening record that identifies the boundary, what was tested and what remains unknown about recall. |
| Systems/discovery | Investigate a product claim or changed result list. | An annotated pipeline or local comparison, with a specific question supported by evidence for the vendor. |
- Information literacy: use the full progression where time allows. If you need to cut, protect Part I, Chapter 11 on query transformation, and Chapters 13–15 on failure, evaluation and professional practice.
- Evidence synthesis: centre Chapters 2 and 11, Chapters 13–15, and Appendix F. Treat Appendix F as core reading because it connects candidate retrieval and recall to active-learning screening and stopping.
- Systems/discovery: emphasise the distinctions in Chapters 4 and 6–11, the diagnostic and governance work in Chapters 13–15, Appendix D for terminology, and Appendix E when products combine or rerank candidate lists.
Three longer course shapes
Chapters have stable links: use the Copy link control in a chapter heading. Each chapter ends with Check yourself questions, with answers hidden until opened. They can serve as reading quizzes; printing reveals the answers. Reading-time estimates exclude lab and exercise work.
A half-day workshop (about three hours)
For practising librarians who want to connect retrieval concepts to a search problem. Use the one-hour workshop as the practical core and add explanations and lab work around it.
- Chapters 1–4 — the whole of Part I. It ends by resolving the category error it exists to correct, so it stands alone better than any other segment.
- Chapter 13 on retrieval failure, which needs surprisingly little of Part II to follow.
- Chapter 15 with Appendix D — the procurement and evaluation material.
A workable 180-minute allocation: 35 minutes on selected Part I explanations; 20 minutes for the BM25 tour and debrief; a 10-minute break; the 60-minute workshop above; 35 minutes adapting its record to a local enquiry or vendor question; and 20 minutes for reporting back and next steps. Treat the chapters listed above as supporting reading, rather than attempting to read all of them during the session. Part II can be assigned later when participants need the fuller account of representations and reranking.
A six-week module
One part a fortnight, with the application exercise as the assessment for each.
| Weeks | Reading | Assessment |
|---|---|---|
| 1–2 | Part I (Chapters 1–4) | Application exercise I |
| 3–4 | Part II (Chapters 5–11) | Application exercise II |
| 5–6 | Part III (Chapters 12–15) | Application exercise III |
Part II covers seven chapters and is the heaviest fortnight. For a shorter module, assign Chapters 5–7 as the core account of embeddings, use Chapters 8–10 to compare pipeline components, and select the query-object examples in Chapter 11 that fit the cohort. Chapter 12 introduces control arrangements before diagnosis and evaluation in Chapters 13–14.
The optional Vector Similarity Lab pairs with Chapter 7. Its guided tour is a short way to establish why cosine similarity ignores vector length, why dot product may not, and why neither score is a relevance judgement; the sandbox can then be left open for students to test their own predictions.
A thirteen-week course
Cover all fifteen chapters in thirteen weeks by pairing Chapters 5–6 in one week and Chapters 9–10 in another. Keep one chapter per week elsewhere. Pair Appendix A (Transformers) with Chapter 6; Appendix B (tokenisation) with Chapters 5–6; Appendix C (index execution) with Chapters 2–3; Appendix D (terminology) with Chapter 15 or keep it open throughout; and Appendix E (learning to rank) with Chapters 9–10. Treat Appendix F as required for evidence-synthesis cohorts alongside Chapters 13–15. Appendix G is an optional application of the pipeline to RAG after Chapter 12.
Use the three application exercises as summative assessments, one per part, and the chapter Check yourself questions for weekly formative checks. Add the product-claim verification task as an assessed piece: assign its eight items across individuals or groups, have them consult current documentation, and ask what changed and how they established it. The resulting record is useful to the next cohort. For evidence-synthesis cohorts, Appendix F’s controlled TAR comparison is a fourth assessable task; check its preparation requirements below before assigning it.
This shape leaves room for the thing the book cannot supply: students running searches in systems your institution actually licenses, and reporting back. Chapters 13, 14 and 15 are considerably more useful to someone who has spent a week failing to make a discovery layer behave.
Using the two labs
Use the BM25 Evidence Lab after Chapter 3, then revisit its admission control after Chapter 4. Ask students to predict the ranking before moving the prescribed slider. The tour fixes other inputs; the sandbox allows independent experiments. Changes to document length also change the collection average.
Budget roughly fifteen minutes for the BM25 tour and ten for the vector tour, plus whatever sandbox time you want; both tours have seven prediction steps. These timings are provisional; allow longer for discussion. The application exercises are the opposite shape — each needs a searching session plus written-up findings, so allow a week rather than a class. Reading-time estimates in the book exclude all of this.
The BM25 tour’s sixth step, The boundary that scoring cannot cross, uses a k slider that cuts records the ranking already scored. Protect time for this step and its debrief: ask which records the next stage receives and whether rescoring that fixed input could recover anything outside it. The seventh step returns to vocabulary and admission rules.
Use the vector lab after Chapter 7. The text descriptions and coordinates are assigned teaching examples, not outputs of an embedding model. Compare direction, magnitude and unit normalisation, then ask why the highest similarity need not satisfy the date and study-design criteria. Students can share a sandbox example by copying its URL.
The application exercises
There are three application exercises, one per part. Prefer an institutional system when local evaluation is the goal. For public access, adapt Exercises I and III to PubMed, OpenAlex or another open search system with suitable features. Exercise II can use public technical documentation plus a recorded behavioural probe. All ask for a written deliverable; access to internal vendor settings is not required.
Before assigning: check that the chosen product supports the required query forms and exposes enough results for the task. Supply a recorded example if access fails. State the product and mode, required output, judgement depth where applicable, and deadline. A short in-class version should use a supplied case and one decision; reserve the full searching and write-up for work between sessions.
Exercise I — Explain a result list without invoking meaning
Assign the full instructions. Output: a one-page account of four observations, plausible explanations and the next evidence needed.
Tests: whether the student can resist reaching for “semantic” as an explanation. They run a natural-language query, record four observations, and must account for each using only query analysis, the execution rule, or query transformation.
A strong answer identifies plausible layers, gives evidence, considers alternatives and marks what remains uncertain. A fully lexical account can be correct; unexplained residue is not a requirement. Do not require students to rule out alternatives the interface cannot distinguish.
Common wrong turn: concluding that a missing query term proves semantic retrieval. That is precisely the inference Chapter 4 dismantles, and it is worth marking as a failure of the exercise rather than a minor error.
Exercise II — Map a product onto the pipeline
Assign the full instructions. Output: a six-stage map labelled by evidence type, with one question for the vendor.
Tests: whether the student can distinguish what a vendor documents from what they inferred. Every stage must be labelled documented, inferred or unknown.
A strong answer justifies each label on the six-stage map. Mark the evidence, not the number of unknown entries. Use documented: not used for an explicitly absent stage. Keep documentation-based claims separate from observations and inferences from the behavioural probe.
Common wrong turn: treating a named model as an architecture. “It uses BERT” fills no stage on the map.
Exercise III — Diagnose a failure, then try to document it
Assign the full instructions. Output: the diagnosis and intervention, a comparison of the two runs, and a record of what could not be established. The one-hour workshop prepares participants for the diagnostic part.
Tests: diagnosis using Chapter 13’s workflow, a measurement from Chapter 14, and the record from Chapter 15.
A strong answer tests a remedy suited to the diagnosis and reports honestly when it did not work. It also explains why each missing piece of the search record matters. Assess that gap list under evidence and explicit limits in the shared rubric below.
Keep the comparison fair: write the relevance criteria first; pool and deduplicate both top tens; judge every pooled record against the same criteria, ideally without knowing its source run; then calculate precision@10 for each. Report recovery of the five known relevant seeds separately. Neither measure establishes complete recall. If fewer than ten results are returned, require the count and denominator convention.
Common wrong turn: treating Chapter 13’s diagnostic lenses as four equivalent categories, or diagnosing every failure as vocabulary mismatch because that is the relationship librarians are most accustomed to seeing.
Appendix F’s controlled TAR comparison — for evidence-synthesis cohorts
Not one of the three application exercises, but the only assessable task aimed squarely at the cohort for whom Appendix F is core reading. It fixes a labelled corpus, varies the feature representation between two active learners, and applies four stopping procedures to the same runs.
Tests: whether the student can separate an effect of the representation from an effect of the stopping rule. Those are different mechanisms and the exercise is built so they can be told apart.
A strong answer attributes recall differences to the stopping rule rather than the feature extractor where the evidence supports that, reports for each simulated stop how many known relevant records were left unseen, and says plainly that a live review cannot inspect that tail without doing the work it hoped to save. The numerical thresholds in the appendix are experimental conditions, not recommendations, and a strong answer says so.
Common wrong turn: reading a WSS@95 improvement as evidence that a live reviewer could have stopped safely. It measures ranking efficiency against a labelled set at a fixed target; it says nothing about knowing when that target has been reached.
Preparation required: Appendix F specifies the experiment; these notes do not supply prepared runs or screening logs. To teach it without asking learners to code, prepare the runs yourself and supply their ordered record identifiers, relevance labels and configurations, with outputs needed for the stopping procedures you assign. Preserve the random-order baseline as well as the two active learners. Have learners compare relevant records found at fixed screening budgets. If you provide only ranking logs, restrict the task to comparisons those logs support; do not ask learners to reconstruct an estimated-recall stopping procedure from them alone.
Worked responses and a short rubric
These examples illustrate reasoning, not observations about a particular live product.
- Exercise I: “The added rare string did not empty the results. It may have been removed during analysis or rewriting, or retained as an optional unmatched term. The visible list cannot distinguish these. A documented analyser or an inspectable processed query would help.” Award credit for the observation, plausible mechanism, alternatives and a discriminating next check.
- Exercise II: “The documentation names lexical candidate retrieval and an embedding rerank, so those stages are documented. It does not name a fusion rule, so fusion is unknown, not proven absent. The same result count after changing a query is an observation, not proof of a fixed controller.” Award credit for source and date, correct stage mapping, justified labels and separation of documentation from experiment.
- Exercise III: “After pooling and judging both top tens, the original has four relevant results and the revised run has six: precision@10 rises from 0.4 to 0.6. Recovery of five seeds rises from two to three. That is encouraging for this need, but does not establish complete recall or general superiority.” Award credit for a testable diagnosis, controlled intervention, complete judgements at the stated depth and an honest record of missing information.
For each exercise, use four equally weighted criteria: accurate observation, sound reasoning, adequate evidence and explicit limits. Do not award marks merely for confidence or for reporting many unknowns.
| Criterion | 0 — absent or misleading | 1 — developing | 2 — demonstrated |
|---|---|---|---|
| Observation | Invents a result or treats a claim as an observation. | Records the main result but omits relevant conditions. | Accurately records the result and conditions, separating observation from documentation. |
| Reasoning | Names a mechanism without support or proposes an unrelated remedy. | Offers a plausible explanation and action but overlooks alternatives. | Considers alternatives and chooses a check or intervention that can distinguish them. |
| Evidence | Supplies no traceable support. | Provides some sources or run details, but the comparison cannot be followed. | Links each claim to the supplied case, dated documentation or recorded run; supports any measurement with its judgements. |
| Limits | Claims certainty or completeness the evidence cannot establish. | Mentions uncertainty without explaining its consequence. | States what remains unknown, why it matters and what further evidence would help. |
For the prepared workshop, the supplied case and revealed trace are the evidence; external research is unnecessary. For live exercises, require dates, modes and source or result identifiers. Give feedback on the weakest criterion and allow a revision after new evidence.
Discussion prompts
After Part I
- Systematic review search strategies are published so they can be inspected and rerun. What exactly is being reproduced — and what is not, given that indexes and analysers change underneath?
- If a database silently made every query term optional tomorrow, how would your users find out? How would you?
- Why does a result missing a typed term not establish semantic retrieval? Use the historical Google example to distinguish strict AND, other admission rules and matching evidence outside the visible page. Which current product would you investigate this way?
After Part II
- Every retrieval-trained embedding space rewards particular proxies for relevance. Whose proxies are in the tools you license, and would they say if asked?
- A reranker cannot rescue a document the retriever never returned. What follows for how you teach searching?
- If a system indexes only abstracts, which of your disciplines suffer most, and can you demonstrate it?
After Part III
- Agency is a trade. Name a task in your service where a fixed workflow is clearly the better buy.
- Your library builds a 40-query evaluation set. Who judges relevance, and what happens when two judges disagree?
- PRISMA-S concedes that some proprietary and similarity-based operations may not be fully replicable. Is that a reporting problem or a procurement one?
- A vendor demonstrates an agentic mode that answers your test question well. What could you ask to find out what it could not have done? This is Chapter 12’s invisible menu, and it is the hardest of its three difficulties to probe from outside.
Before you teach it: what will have moved
Product descriptions age faster than the mechanisms they illustrate. Re-check the examples you intend to use; the prepared one-hour case is explicitly hypothetical and requires no current product claim.
Use the table below to choose claims to verify. In a copy, complete Documented as of and Last checked, then attach the source URL, product mode, relevant passage and your finding: unchanged, changed or not established. A recent check of an old document does not establish that the architecture is current. If it cannot be verified, present it as a dated example. Pass that record to the next instructor.
| Where | Product or claim | Documented as of | Last checked | What would change it |
|---|---|---|---|---|
| Chapter 3 and its infrastructure footnotes | Scholarly systems on Lucene-family infrastructure | — | A vendor migrating its search stack | |
| Table 7.2 | Semantic Scholar and OpenAlex Alice | — | Either could change route without changing the interface | |
| Table 9.2, Figure 9.5 | Primo Research Assistant, traced end to end | — | Documented cut-offs. The book's most load-bearing product example: if the numbers move the argument holds but the figures do not | |
| Chapter 10's examples | Web of Science Smart Search (blend), Scopus AI (route) | — | Whether a product blends or routes, which a release note changes quietly | |
| Tables 11.4, 11.5 | Query transformation and filter extraction across products | — | Vendors add supported fields steadily | |
| Table 12.3 | Six academic tools by control arrangement | April 2026 | A vendor renaming or re-scoping a mode. The same vendor already appears on two rows, so brand-level generalisations were wrong even then | |
| Appendix G | Scopus AI and RAG Fusion walkthrough | 2024 | Already labelled historical. Check before presenting it as anything but a worked example of reading documentation | |
| Figure 1.3 | The two-by-two of AI academic search | — | Products migrate between quadrants |
Handing these rows out one per student is also the assessed task suggested for the thirteen-week course above. It dates the material for you, and the filled-in table is worth more to the next cohort than another essay would be.
Figures, and reusing them in slides
The book’s original text, figures and these teaching materials are licensed CC BY 4.0. You may excerpt, adapt, translate and redistribute them, including commercially, with credit and without asking. Third-party screenshots and reproduced research figures, including Figure 5.5 from Mikolov and colleagues, retain their original rights and attribution.
The original explanatory illustrations live in images/ in the repository and are CC BY 4.0 along with the rest of the text — use them in slides with credit. They were generated with AI assistance, as the book's disclosure records.
The numbered figures built as HTML rather than images — the BM25 scoring walkthrough, the pruning panels, the WordPiece pipeline — are in the page itself and are easiest to capture by screenshot at a wide window. Each has a stable anchor, so you can also just link to it.
Commercial screenshots and reproduced research figures are not covered by the book’s licence. They are third-party material reproduced for comment and criticism. If you are redistributing adapted material, take your own or leave them out.
Reference tools for learners and instructors
The glossary is the most directly assignable thing in the book for an information-literacy cohort: around a hundred short definitions, each written to separate a mechanism from the label usually attached to it. It sits in a collapsed panel in the end matter, so students who never open it will not find it. Assigning twenty terms and asking which ones name a representation, which a pipeline stage and which only an interface is a reasonable first seminar. Appendix D is the same move at greater length, and reads well alongside it.
The changelog records what changed between versions. If you assigned a chapter last term and it reads differently this term, that is where to look — and it is what makes the version number in the citation below worth recording.
Corrections are welcome through the repository. Include the affected passage, product and mode, source URL and date, and what the current documentation establishes.
Leave a teaching record for the next instructor
After the first delivery, record the audience, number of participants, material used, preparation time and actual session length. Note where learners needed help, how their initial explanations changed, and whether they could handle the transfer case independently. Use anonymised examples and omit confidential enquiry details.
Ask participants a week or two later whether they used the method in a consultation, class or evaluation, and what helped or prevented that use. Record any actual reuse of the session by a colleague. These observations help distinguish an enjoyable session from a resource that changes practice.
Copyable record: audience and date; preparation minutes; session minutes; most difficult distinction; evidence from the transfer check; one change for next time; later use reported. Treat the proposed timings as provisional until you have this evidence.
Citing it
Cite the version as well as the date. The page is live and will change; the version number is what lets a student say afterwards which text they read.
Tay, A. C. H. (2026). How search decides what you see: A librarian’s guide to Boolean search, BM25, embeddings, reranking, and the retrieval pipelines behind hybrid and agentic search (Version 1.2.1). https://aarontaycheehsien.github.io/Information-retrieval-crashcourse/