03 — Making content memorable: elaboration, imagery, voice¶
Where 01 says how often to touch material and 02 says how much per breath, this section says
what the extra words should be.
3.1 Dual coding and the concreteness effect — [A] for the effect, [B] for instructed imagery¶
Paivio's dual coding theory: verbal and imagistic representations are separate but linked stores, and material that lands in both is recalled better. The concreteness effect — concrete words are remembered better than abstract ones — is one of the most replicated findings in memory research, across languages, ages, recall and recognition.
Crucially for an audio-only product: you do not need a picture to get the imagery benefit. Imagery instructions ("picture a ...") and concrete language both recruit the image store. What matters is that the image is actually generated by the learner, so imagery prompts need a moment of silence to work.
There is a machine-usable resource here: Brysbaert, Warriner & Kuperman (2014) concreteness ratings for ~40 000 English lemmas (1–5 scale), plus age-of-acquisition and prevalence norms. These give us a cheap, deterministic abstractness signal per passage — no LLM required.
Implications.
- Score every passage for abstractness using concreteness norms; abstract-and-important passages
earn a concrete anchor: a short, vivid, physically imaginable scene that instantiates the
concept, followed by a brief pause to let the listener build it.
- An anchor must be stable: the same concept gets the same anchor every time it recurs in the
document, so the anchor itself becomes a retrieval cue. This requires a persistent concept
registry, not per-paragraph improvisation.
- Prefer concrete vocabulary in generated scaffolding, even when the source is abstract.
→ Rules IMG-*.
3.2 Elaborative interrogation and self-explanation — [B] moderate utility¶
Dunlosky et al. (2013) rate both as moderate utility: real, generalizable benefits, held back from "high utility" mainly by a thinner classroom evidence base. Elaborative interrogation = generating why a stated fact is true. Self-explanation = explaining how new information relates to what you already know. Even brief elaboration substantially raised recall of fact sets.
Implications.
- Facts should rarely be delivered bare. For each retained key claim, the writeup answers why is
this true / why does it follow in one or two sentences — this is generated content, so it must be
grounded in the source and flagged when it is not (see the groundedness gate in the plan).
- Where the listener can plausibly do the work, convert the elaboration into a prompt-then-answer
pair instead of a statement — that gets the generation effect too.
→ Rules ELB-*.
3.3 Analogy and prior knowledge¶
Comprehension is a function of what the listener already knows. Analogy is the standard bridge, and it is also the standard failure mode: a wrong analogy installs a durable misconception. Anchors and analogies must be marked as ours, not the author's, so that the listener can discount them.
Implications. Every generated analogy is explicitly framed ("Here is an analogy — it is mine,
not the paper's, and it breaks down when ..."). Stating the limits of the analogy is not optional;
it is what stops the analogy from becoming the memory. → Rules ANA-*.
3.4 Narrative superiority — [A] for the effect, [C] for our application of it¶
A meta-analysis of 150 effect sizes / >33 000 participants found memory and comprehension of narrative text is superior to expository text ("stories are better recalled than essays"), largely because narrative structure is highly predictable and coherence is easy to maintain.
But the neighbouring literature warns us: informative narratives raised situational interest without raising comprehension in a 2024 study, and seductive details — interesting but tangential material — are recalled at the expense of main ideas.
Implications. Use narrative structure (a problem, an attempt, an obstacle, a resolution — which
maps neatly onto IMRaD: why anyone cared, what they tried, what got in the way, what it changed) as
the organizing frame for sections. Do not add narrative decoration: no invented anecdotes, no
colourful asides, no "fun facts". Story shape, not story flavour. → Rules NAR-*, COH-*.
3.5 Personalization and voice — [B]¶
Mayer's personalization principle: conversational, second-person wording beats formal wording, with a median d ≈ 1.11 across 11 of 11 comparisons (with caveats: weaker for high-achieving learners and long lessons). The voice principle: a human(-like) voice beats a machine voice (median d ≈ 0.74, 5 of 6 comparisons) — modern neural TTS plausibly narrows this, but it argues for spending effort on voice quality and natural phrasing.
Implications. All generated scaffolding — orientation, prompts, recaps, glosses — is written in
direct second person ("you"), in a spoken register, with contractions. Source material keeps its
own claims and precision; we change the frame, not the facts. Because the personalization
effect weakens over long lessons, this is a style default, not a licence for chattiness — the
coherence principle outranks it. → Rules VOI-*.