Skip to content

03 — Making content memorable: elaboration, imagery, voice

Where 01 says how often to touch material and 02 says how much per breath, this section says what the extra words should be.


3.1 Dual coding and the concreteness effect — [A] for the effect, [B] for instructed imagery

Paivio's dual coding theory: verbal and imagistic representations are separate but linked stores, and material that lands in both is recalled better. The concreteness effect — concrete words are remembered better than abstract ones — is one of the most replicated findings in memory research, across languages, ages, recall and recognition.

Crucially for an audio-only product: you do not need a picture to get the imagery benefit. Imagery instructions ("picture a ...") and concrete language both recruit the image store. What matters is that the image is actually generated by the learner, so imagery prompts need a moment of silence to work.

There is a machine-usable resource here: Brysbaert, Warriner & Kuperman (2014) concreteness ratings for ~40 000 English lemmas (1–5 scale), plus age-of-acquisition and prevalence norms. These give us a cheap, deterministic abstractness signal per passage — no LLM required.

Implications. - Score every passage for abstractness using concreteness norms; abstract-and-important passages earn a concrete anchor: a short, vivid, physically imaginable scene that instantiates the concept, followed by a brief pause to let the listener build it. - An anchor must be stable: the same concept gets the same anchor every time it recurs in the document, so the anchor itself becomes a retrieval cue. This requires a persistent concept registry, not per-paragraph improvisation. - Prefer concrete vocabulary in generated scaffolding, even when the source is abstract. → Rules IMG-*.

3.2 Elaborative interrogation and self-explanation — [B] moderate utility

Dunlosky et al. (2013) rate both as moderate utility: real, generalizable benefits, held back from "high utility" mainly by a thinner classroom evidence base. Elaborative interrogation = generating why a stated fact is true. Self-explanation = explaining how new information relates to what you already know. Even brief elaboration substantially raised recall of fact sets.

Implications. - Facts should rarely be delivered bare. For each retained key claim, the writeup answers why is this true / why does it follow in one or two sentences — this is generated content, so it must be grounded in the source and flagged when it is not (see the groundedness gate in the plan). - Where the listener can plausibly do the work, convert the elaboration into a prompt-then-answer pair instead of a statement — that gets the generation effect too. → Rules ELB-*.

3.3 Analogy and prior knowledge

Comprehension is a function of what the listener already knows. Analogy is the standard bridge, and it is also the standard failure mode: a wrong analogy installs a durable misconception. Anchors and analogies must be marked as ours, not the author's, so that the listener can discount them.

Implications. Every generated analogy is explicitly framed ("Here is an analogy — it is mine, not the paper's, and it breaks down when ..."). Stating the limits of the analogy is not optional; it is what stops the analogy from becoming the memory. → Rules ANA-*.

3.4 Narrative superiority — [A] for the effect, [C] for our application of it

A meta-analysis of 150 effect sizes / >33 000 participants found memory and comprehension of narrative text is superior to expository text ("stories are better recalled than essays"), largely because narrative structure is highly predictable and coherence is easy to maintain.

But the neighbouring literature warns us: informative narratives raised situational interest without raising comprehension in a 2024 study, and seductive details — interesting but tangential material — are recalled at the expense of main ideas.

Implications. Use narrative structure (a problem, an attempt, an obstacle, a resolution — which maps neatly onto IMRaD: why anyone cared, what they tried, what got in the way, what it changed) as the organizing frame for sections. Do not add narrative decoration: no invented anecdotes, no colourful asides, no "fun facts". Story shape, not story flavour. → Rules NAR-*, COH-*.

3.5 Personalization and voice — [B]

Mayer's personalization principle: conversational, second-person wording beats formal wording, with a median d ≈ 1.11 across 11 of 11 comparisons (with caveats: weaker for high-achieving learners and long lessons). The voice principle: a human(-like) voice beats a machine voice (median d ≈ 0.74, 5 of 6 comparisons) — modern neural TTS plausibly narrows this, but it argues for spending effort on voice quality and natural phrasing.

Implications. All generated scaffolding — orientation, prompts, recaps, glosses — is written in direct second person ("you"), in a spoken register, with contractions. Source material keeps its own claims and precision; we change the frame, not the facts. Because the personalization effect weakens over long lessons, this is a style default, not a licence for chattiness — the coherence principle outranks it. → Rules VOI-*.