Elena: Welcome to Emma’s Library Research Dialogues. I’m Elena. Eli: And I’m Eli. Today’s paper asks whether two synthetic voices can teach better than one, which means we are either the ideal hosts or a serious conflict of interest. Elena: We are a demonstration, not an experiment. Eli: That was immediate. Elena: It is the distinction the paper needs. The title of our episode is Can Two AI Voices Teach Better Than One? The source is a 2026 education-technology paper by Gendo Kumoi and colleagues. They built a human-reviewed system for turning lesson material into either a single synthetic narration or an expert-and-novice dialogue, then studied the formats with 245 first-year high-school students. Eli: And the result people will remember is that dialogue won. Elena: Won what? Eli: Enjoyment, very clearly. Among the students who answered that preference question, 66.9 percent chose dialogue as the most enjoyable of the three video formats. Students in the dialogue session also reported stronger understanding, more confidence that they could explain the lesson, and more active thinking on several measures. Elena: All true. Now put the caveat beside it, not five minutes later. Eli: The formats came in a fixed order, on different dates, with different lesson content. The outcomes were questionnaires and preferences, not objective tests of what students learned or retained. So the study found a promising pattern. It did not prove that dialogue caused better learning. Elena: Good. That is our central tension today. Dialogue may help a learner notice a question, hear reasoning unfold, and test a mental model. It can also sound less natural, add switching costs, or create the feeling of understanding without the durable knowledge underneath. Eli: I like the format because it makes the thinking audible. A single narrator can give me a polished explanation, but a second voice can interrupt at exactly the point where I quietly stopped following. Elena: If the second voice asks the question you actually have. In the study, only 39.8 percent agreed that the novice’s questions resembled their own. That is not failure, but it is a warning against assuming that one scripted novice represents every learner. Eli: So today we will do three things. We will explain the learning theory behind expert-and-novice dialogue, open the system the researchers built, and read the classroom evidence without asking it to carry more weight than it can. Elena: And because Emma’s Library already has a single-narrator review of this paper, listeners can compare the forms directly. Same source, different experience. Eli: Although not a controlled comparison. Elena: You are learning. Eli: I resent how satisfying that was. Let’s begin with why a second voice might matter at all. Eli: Imagine one voice saying, Here is the concept, here is the definition, and here is the example. Efficient. Now imagine a second voice saying, Wait—why does the example fit? The lesson has not gained new information yet. It has gained a visible place to think. Elena: That second voice can perform several jobs. It can surface a likely misconception, request a concrete case, try a rephrasing, or ask what changes under a different condition. The learner hears not only an answer but a way of approaching uncertainty. Eli: The paper connects that to vicarious learning. You can learn by observing someone else engage with an explanation, especially when the observed learner asks, attempts, and revises instead of merely receiving the right answer. Elena: It also draws inspiration from cognitive apprenticeship. Traditional apprenticeship makes expert practice visible through modeling, coaching, and scaffolding. In a dialogue lesson, the expert can verbalize a reasoning process; the novice can expose where support is needed; the explanation can then become more structured. Eli: The researchers turned that into a five-stage pattern. The expert introduces a concept. The novice asks a question. The expert elaborates. The novice rephrases. The expert confirms, corrects, or extends. Elena: We should demonstrate it with something ordinary. Eli: All right. Spaced practice improves retention because retrieving an idea after some forgetting strengthens later access more than repeating it immediately. Elena: Why would partial forgetting help? That sounds like losing ground. Eli: Because the effort of reconstructing the idea is part of the learning event. If the answer is still sitting in working memory, another repetition may feel fluent without requiring retrieval. After a delay, successful recall exercises the route back to the idea. Elena: So the gap is not valuable because forgetting is good. It is valuable because the later attempt makes me recover the idea rather than merely recognize it. Eli: Yes, with one correction: the delay has to remain manageable. If retrieval fails completely and there is no feedback, difficulty can become discouragement rather than productive effort. Elena: That small correction is the interesting part. A perfect rephrasing followed by yes can become ceremonial. The learner-shaped voice needs permission to misunderstand in a useful way. Eli: And the expert has to show reasoning, not just award marks. The paper suggests question types such as why, how, what-if, association, and confirmation. Those are not decorations. They create different cognitive moves. Elena: Give me the difference between confirmation and what-if. Eli: Confirmation checks the current model: So the delay matters because I must retrieve? A what-if tests its boundary: What if I wait so long that I remember nothing? One stabilizes an idea; the other tries to break it. Elena: This is where dialogue can be more than alternating sentences. One host may hold the floor long enough to build a model. The other listens, identifies the unstable part, and returns with a question that changes the explanation. Eli: Which is not the same as permanent expert and permanent novice. Elena: I prefer rotating roles. Expertise depends on the question, and adults do not become more believable because one is scripted to be confused forever. Eli: The paper’s production system did use explicit Expert and Novice roles, with roughly a seven-to-three speech balance. That makes sense for a lesson. For a continuing podcast, the deeper principle is temporary asymmetry: one person models; the other interrogates, rephrases, or transfers. Elena: The educational value is not two voices. It is two different cognitive functions made audible. Elena: The phrase AI-generated lesson can make the process sound like one button. The system in this paper is deliberately not that. Eli: It has three broad stages. First, an educator supplies the lesson material and requirements, and the system uses a language model to create slides and a narration script. Second, the educator reviews and corrects the content and the instructional design. Third, the system synthesizes speech and integrates the audio with the slides. Elena: Human review occurs inside the production loop, not as a hopeful glance after publication. Eli: There are two substantial feedback points. The educator checks the generated slides for accuracy, structure, readability, and visual fit. Then the educator reviews the narration for clarity, factual correctness, difficulty, and instructional style. Elena: Why not ask the model to judge its own work and remove the human bottleneck? Eli: Because the educator knows the learners, the curriculum, and the cost of a mistake in that setting. A model can help identify inconsistencies, but it does not own the teaching responsibility. The authors explicitly describe the system as semi-automated and say it is not intended to replace educators. Elena: I agree with the responsibility point. I am less convinced by the word bottleneck. Sometimes the narrow place is where quality is made. Eli: Fair. But production cost still matters. If synthetic speech removes the need for an instructor to record every revision, a teacher may be able to make more material or update it faster. Elena: Provided the saved recording time is not quietly replaced by hours repairing scripts and unstable voices. Eli: Which is why the paper is useful as a system paper, not only a format comparison. It gives concrete TTS writing rules: calibrate the audience and vocabulary, control sentence length, use punctuation for pacing, disambiguate symbols and difficult readings, and write emotion in language the speech engine can express clearly. Elena: The last point deserves attention. A human speaker can rescue a stiff sentence with timing, facial expression, or a glance at the room. Synthetic speech needs more of that intention encoded in the text. Eli: The system also offers two narration modes. Single-speaker mode follows a structured lecture: opening, introduction, main content, reinforcement, confirmation, summary, closing. Dialogue mode assigns an expert and novice, then applies the five-stage learning pattern we just used. Elena: And it generates audio in segments aligned with slides. That is practical for editing, but it later created a quality problem: prosody and voice identity could vary from one generation to the next. Eli: Especially in dialogue, where listeners are tracking two people. A small shift that sounds like harmless variation in one narrator can make a two-speaker scene feel as if one host changed identity between slides. Elena: Our own pipeline learned that lesson. We render manageable scenes, preserve stable voice assignments, check speaker identity, and listen across boundaries rather than judging isolated clips. Eli: That is a design implication from the paper, not evidence that our version teaches better. Elena: Correct. We have improved the production conditions. We have not run the learning study. Eli: There is another choice in the paper I like: the novice is written to ask learner-shaped questions, not merely provide comic relief or emotional reactions. Elena: Yet the classroom results suggest that function was only partly achieved. Fifty-seven point three percent said the novice’s questions helped them understand. Only 39.8 percent said those questions resembled their own. Eli: So helpful without always feeling like me. Elena: Yes. A scripted learner can scaffold a concept without successfully simulating the listener’s mind. Designers should not confuse those achievements. Eli: Now the study. Two hundred forty-five first-year high-school students experienced three video lessons as part of an inquiry-based learning program. Elena: In the same sequence for everyone. Eli: Yes. First came a traditional lesson video with an instructor’s human voice. About two months later came a different lesson with single-speaker synthetic narration. About one month after that came another lesson using synthetic expert-and-novice dialogue. Elena: Which means format, content, date, and order changed together. Eli: They did. The authors state clearly that they cannot isolate a causal format effect. Their goal is exploratory: describe patterns and derive design implications. Elena: I want to make the confound intuitive. Suppose the dialogue lesson receives better ratings. Perhaps dialogue helped. Perhaps the topic was more interesting. Perhaps students were more familiar with the course by the third session. Perhaps the school calendar changed their attention. The study cannot fully separate those stories. Eli: There was also a prior-knowledge imbalance. Sixty-six percent of respondents in the dialogue questionnaire said they knew nothing at all about that lesson’s content, compared with 34.1 percent in the single-TTS questionnaire. Elena: So the dialogue session began at a disadvantage on measured prior knowledge. Eli: Yes. The authors treated that imbalance as a confound and added a proportional-odds model using prior knowledge as a covariate. They argue that the unadjusted dialogue advantages are conservative in that respect. Elena: Though one covariate does not erase the other differences in content, order, timing, or who completed each questionnaire. Eli: After the single-voice lesson, 229 questionnaires were valid. After the dialogue lesson, 206 were valid. The questionnaires were anonymous, so even though the same school cohort likely contributed to both, individual responses could not be matched. Elena: That matters because the detailed comparison treats the two response sets approximately as independent. If many are the same learners, their answers are correlated. If the responders differ, selection may also matter. Eli: The paper therefore calls it a repeated cross-sectional comparison, not a clean paired experiment. It uses Mann–Whitney tests for detailed items, applies Benjamini–Hochberg false-discovery-rate correction across twenty comparisons, and adds the prior-knowledge model as a supplementary analysis. Elena: Listeners do not need to memorize those names. They need to know what the safeguards were trying to prevent. Eli: The rank-based test compared the distributions of ratings. The false-discovery correction reduced the chance that one apparently significant item emerged simply because the authors tested many. The proportional-odds model asked whether key differences remained after accounting for measured prior knowledge. Elena: Useful safeguards inside an imperfect design. Statistical care can make an estimate more responsible; it cannot randomize yesterday. Eli: That is going on a mug. Elena: It will sell terribly. Eli: The study also asked students, after the final session, to retrospectively rate the three formats on comprehension, concentration, and overall evaluation. For those comparisons, 182 or 183 students had complete responses. Elena: But the human-voice lesson was roughly two months in the past and the single-TTS lesson roughly one month in the past. Retrospective ratings can smooth or distort experience with time. Eli: And once again, each format carried different content. So when we reach the numbers, we have to say what was actually compared: remembered experiences of three different sessions, not three audio tracks laid over the same lesson. Elena: This is why I resist the headline two voices teach better. The design can show that the complete dialogue session was received differently. It cannot tell us how much of that difference belongs to the second voice. Eli: I still think the pattern is useful. Elena: It is. A pattern can guide the next experiment without pretending to be the last word. Elena: The first research question was whether synthetic speech substantially degraded the learning experience compared with the instructor’s voice. Eli: On the three retrospective core ratings—ease of comprehension, ease of concentration, and overall evaluation—the means were very close. The Friedman tests found no statistically significant differences among the human instructor, single TTS, and dialogue TTS sessions. Elena: No significant difference does not prove equivalence. Eli: The authors say that too. They added two one-sided equivalence tests, usually called TOST, with a margin of plus or minus half a point on the five-point scale. All pairwise differences for the three core measures fell within that range, with p-values below point zero zero zero one. Elena: Translate the claim exactly. Eli: The observed retrospective differences were small enough to fit inside the study’s chosen half-point equivalence range. That gives more support than a non-significant test alone for saying the sessions did not look substantially different on those ratings. Elena: But it does not establish that human and synthetic voices are educationally interchangeable. Eli: Because the content and dates differed, the ratings were retrospective, and the half-point margin was a conventional ten percent of the scale rather than a demonstrated minimum educationally meaningful difference. Elena: Good. I am pressing this because equivalence language is unusually easy to overstate. The study supports a limited practical suggestion: in this setting, synthetic narration did not produce clearly worse remembered ratings on those three measures. Eli: Which matters. A teacher who dislikes recording, needs rapid revisions, or wants several character roles may gain production flexibility without an obvious collapse in the reported experience. Elena: May. The paper did not measure total production labor, cost across subjects, accessibility outcomes, or what happens when a less technical educator runs the system. Eli: You have a remarkable ability to put seatbelts on a sentence. Elena: And you have a remarkable ability to drive it before the doors close. Eli: Fair. Still, this result shifts the practical question. Instead of asking, can synthetic speech ever be acceptable, we can ask, under which learning goals is a single narrator better, and when does dialogue earn its added complexity? Elena: That is the useful transition. The average core ratings did not separate the formats. The more detailed items reveal a trade-off rather than a universal winner. Eli: On self-assessed understanding, 78.2 percent of dialogue respondents gave a high rating, compared with 69 percent after single TTS. For I can explain the main content to a friend, the figures were 43.7 percent for dialogue and 31 percent for the single narrator. Elena: Those differences remained statistically significant after false-discovery correction, with small effect sizes. In the supplementary model controlling for prior knowledge, the direction also remained: the estimated odds of a higher response were greater for dialogue on understood the lesson and can explain to a friend. Eli: Dialogue respondents were also more likely to say the time was worthwhile and that they tried to deepen their thinking. Then there is the preference result: 66.9 percent selected dialogue as the most enjoyable video format, far above the instructor video and single TTS. Elena: Yet the detailed item this lesson was enjoyable did not remain significant under the main corrected rank-based comparison. It became significant only in the supplementary prior-knowledge model. Eli: Does that make the 66.9 percent preference meaningless? Elena: No. It means preference and an absolute rating answer different questions. A learner can choose dialogue as the most enjoyable of three options without giving the dialogue lesson an exceptionally high enjoyment score. And the preference sample was 154 respondents, not all 245. Eli: I still hear a coherent signal: dialogue helped some students organize the material and was the format many wanted more of. Elena: I hear that too. I refuse the word helped if it quietly means increased objective learning. The outcomes were immediate self-reports. No retention test, transfer task, exam score, or later viewing behavior was measured. Eli: Then say supported confidence and engagement. Elena: Better, with the design caveats still attached. Eli: Now the cost. Single-speaker TTS sounded more natural. Fifty-five point five percent rated its audio natural, compared with 35.4 percent for dialogue. The difference was highly significant, with a small-to-medium effect. Single TTS was also rated easier to hear. Elena: That result is not cosmetic. Following unstable speaker cues consumes attention. If the voice identity or prosody shifts between segments, the listener spends cognitive effort working out who is speaking instead of what the explanation means. Eli: The authors describe that as a pattern consistent with additional extraneous cognitive load. Yet students still reported stronger confidence and active thinking in dialogue. That is the trade-off I find exciting: less polished sound, more visible cognitive structure. Elena: Or more novelty, different content, later-session familiarity, and a structure that feels engaging. The paper cannot allocate the effect among those possibilities. Eli: You are not letting me have even one clean sentence. Elena: The evidence did not give us one. Eli: All right. Here is the less clean version. In this exploratory study, the dialogue session produced several encouraging self-report signals despite being rated less natural. That justifies improving the audio and running a stronger learning experiment. Elena: Yes. And the novice result tells us where to improve the script. Fifty-seven point three percent found the novice’s questions helpful, but only 39.8 percent felt those questions were like their own. The novice worked better as cognitive scaffolding than as emotional identification. Eli: Meaning we should prioritize questions that expose relationships, assumptions, and boundaries rather than trying too hard to make the novice sound adorably relatable. Elena: Exactly. A learner proxy does not need to impersonate the audience. It needs to perform a useful act of learning in public. Eli: That might be the most actionable line in the paper. Elena: If you were designing the next study, what would you change first? Eli: Same lesson content in every format. Random assignment to single narrator or dialogue. Enough learners to detect a practically meaningful difference. Then an immediate knowledge test, a delayed retention test, and at least one transfer problem that looks different from the examples. Elena: I would add a crossover design if practical, with counterbalanced order, so some learners hear dialogue first and others single TTS first. I would retain anonymous privacy protections while giving each participant a random identifier, allowing paired analysis without collecting names. Eli: I would measure production too: educator review time, number of corrections, regeneration failures, audio defects, and cost per finished minute. A format that gains a small engagement advantage but doubles the hidden labor may not scale. Elena: Accessibility should not be an afterthought. Some learners may benefit from distinct voices. Others may find switching harder, especially when voice identity is unstable. Subtitles, speaker names, and consistent visual cues could reduce the burden. Eli: The paper recommends those cues. It also suggests improving prosody control and adapting the novice’s question strategies and amount of speech to the learner and subject. Elena: Adaptation creates another study. If the system chooses questions based on a learner profile, we need to know whether the profile is accurate, whether students can correct it, and whether personalization narrows the material too aggressively. Eli: You have brought agent memory into the dialogue-TTS episode. Elena: We said the hosts could remember previous work. Eli: That is either continuity or a crossover event. Elena: Keep moving. Eli: For actual production today, I would not choose one format for an entire course. Use single narration when the material is linear, listening ease is the priority, or the explanation does not benefit from visible questioning. Use dialogue for conceptual thresholds: places where learners predictably confuse cause and correlation, procedure and principle, evidence and interpretation. Elena: I would be even more selective. Dialogue should earn each switch. A new voice should ask a question the lesson needs, test an idea, or carry a different perspective. If the second speaker merely repeats the first in friendlier words, the listener pays a switching cost for little gain. Eli: Our comparison in Emma’s Library can make that tangible. The single-narrator version is calmer, denser, and more uniform. This version exposes our uncertainty and makes the evidence negotiation audible. Elena: It also takes longer and includes our judgments about what deserves disagreement. Listeners may learn more, less, or simply differently. Completion rate would tell us what people stayed with; it would not by itself tell us what they understood. Eli: We could offer the same five-question quiz after both versions. Elena: Useful for a product comparison if assignment is handled carefully. But voluntary listeners choose their format, and repeat listeners have already encountered the material. We should not turn ordinary library analytics into a causal study by changing the label. Eli: There is the seatbelt again. Elena: The doors are still open. Eli: I do think the comparison has value even without a verdict. It lets a listener notice personal friction: Did the questions help? Did the voices make structure clearer? Did the banter distract? Which claims can you explain afterward? Elena: That last question is the one I care about. Delight is not the enemy of learning. It becomes a problem only when we use the feeling of fluency as a substitute for checking what remains. Eli: Then the best podcast pipeline is not the one that maximizes realism. It is the one that uses human texture in service of attention, explanation, and recall. Elena: With enough realism that people want to stay. Eli: I will take the concession. Elena: It was a specification. Eli: Let’s leave the listener with the paper’s strongest lesson and its strongest limitation. Elena: The strongest lesson is that instructional dialogue is a design pattern, not merely a voice count. An expert can model reasoning. A novice can ask, test, rephrase, and expose a boundary. Human review connects those moves to real learners and real subject matter. Eli: The strongest empirical signal is that, in this study, the dialogue session was preferred for enjoyment and received better ratings on several measures of confidence and active engagement, even though the single synthetic narrator sounded more natural. Elena: The strongest limitation is that fixed order and changing lesson content prevent a causal claim about format. The outcomes were self-reports, the detailed questionnaires could not be matched person by person, the rating scales were created for the study, and long-term learning was not tested. Eli: So can two AI voices teach better than one? Elena: Sometimes they may create a better opportunity to learn. This paper does not yet show that they produce more learning. Eli: I would say two voices are worth trying when the second one performs real intellectual work. Elena: And worth removing when it does not. Eli: If you want to compare the experience yourself, Emma’s Library has both versions: the original single-narrator audio review and this Elena-and-Eli research dialogue. Listen to either one, or try the same section in both. Elena: Then, without replaying, answer three questions. What did the classroom study directly measure? What did dialogue appear to improve? And what design choice stops us from saying it caused better learning? Eli: If you can explain those to someone else, the format has at least helped you organize the evidence. Elena: Or you did the work. Eli: I was including the listener in the format. Elena: I know. I am protecting their credit. Eli: Fair enough. The source was A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential, by Gendo Kumoi and colleagues, version one of arXiv two six zero seven point one two two three five. Elena: It was a 2026 preprint when this episode was prepared, so later peer review or revisions may change the record. Our discussion is independent educational commentary, not a claim that synthetic dialogue should replace teachers. Eli: I’m Eli. Elena: And I’m Elena. Thanks for listening to Emma’s Library Research Dialogues. Eli: We’ll meet you in the next question.