# Can Two AI Voices Teach Better Than One? One Paper a Day. Today: a 2026 education-technology preprint about synthetic speech, expert-and-novice dialogue, and the difference between a promising classroom pattern and proof that students learned more. This is independent educational commentary on version one of arXiv paper 2607.12235. The manuscript is a preprint submitted for consideration to IEEE Access and had not been peer reviewed when this episode was prepared. ## 1. The question and the bottom line Here is the finding first. In a practical study with 245 first-year high-school students, a lesson narrated as an expert-and-novice synthetic dialogue received better ratings than a single synthetic narrator on self-assessed understanding, ability to explain the lesson, and active thinking. Dialogue was also the clear favorite for enjoyment: 66.9 percent of respondents who answered that preference question selected it. The biggest caveat belongs in the same breath. The three formats were presented in a fixed order, on different dates, with different lesson content. The outcomes were questionnaires and preferences, not objective tests of what students learned or remembered. The study therefore does not prove that dialogue caused better learning. Who should care? Teachers deciding whether synthetic speech can make lesson production more practical; learning-product teams choosing between a clean single narrator and a more conversational format; and anyone evaluating claims about artificial intelligence in education. The responsible conclusion is narrow but useful. Human-reviewed dialogue TTS looks promising for engagement and confidence, while single-speaker TTS sounded more natural. The next experiment must keep the lesson constant, randomize format and order, and measure actual learning. ## 2. Why two voices might matter A normal synthetic lecture asks one voice to carry everything: introduce a concept, explain it, anticipate confusion, and summarize it. A dialogue can distribute those jobs. An expert explains. A novice interrupts with the question a learner might be forming, tries a restatement, and receives correction. This design has a plausible learning mechanism. In a 2008 Cognitive Science study, Chi and colleagues found that pairs of students who collaboratively observed a recorded human tutorial learned to solve physics problems as effectively as individually tutored students in that experiment. The relevant idea is vicarious learning: an observer can learn by watching someone else struggle, ask, and refine an explanation. But that evidence involved human tutoring and collaboration. It supports the theory behind a novice voice; it does not automatically validate an artificial dialogue. There is also a competing mechanism: cognitive load. Working memory has limited capacity for unfamiliar information. A second voice may provide structure, or it may add identification work, awkward transitions, and distracting variation. The updated review of cognitive load theory by Sweller, van Merrienboer, and Paas is useful here because it turns the design question into a trade-off. Does dialogue help learners organize ideas enough to justify the extra audio complexity? Earlier classroom evidence gives a reason not to assume that modern synthetic speech is harmless. A 2022 K-12 study by Dai and colleagues compared Dutch text-to-speech models and an original human voice. The human voice performed better on listening experience and knowledge-test scores. That result does not settle the present question, but it warns that voice quality and comprehension must be tested rather than treated as solved. ## 3. The system the researchers built Gendo Kumoi and six coauthors designed a three-stage, human-in-the-loop production system. The phrase human-in-the-loop matters. Their proposal is not to press one button and trust whatever the model creates. In stage one, an LLM turns textbook or lecture material into Markdown slides using Marp. The prompt specifies structure, headings, visual elements, and layout. An educator then fact-checks the slides, repairs the instructional flow, and adds or changes figures and tables. In stage two, the system writes narration for each slide. The authors argue that prose written for the page can sound stiff when read aloud, so the prompt controls audience level, vocabulary, sentence length, pauses, rhythm, pronunciation, and coordination with visuals. The educator reviews again for accuracy, clarity, and teaching style. In stage three, a speech model generates the audio. The system combines it with rendered slides to produce a lesson video. The paper used a Gemini text-to-speech preview model available at the time of the study. The implementation inserted silence around segments and used prompt tags to shape pauses and intonation. The workflow's strongest design choice is not the model name. Models change. It is the decision to place an educator after content generation and again after script generation. That makes factual and pedagogical judgment an explicit production stage rather than an invisible hope. What the paper does not yet quantify is equally important: how much teacher time the workflow saves, how many corrections each stage requires, and which kinds of errors survive review. The authors name that measurement as future work. ## 4. How the expert-and-novice script works The dialogue mode assigns roughly seventy percent of the speech to the expert and thirty percent to the novice. A typical sequence has five moves. The expert introduces a concept. The novice asks a question. The expert elaborates. The novice restates the idea. The expert confirms and extends it. The novice can ask why, ask how, explore a hypothetical, connect ideas, or check understanding. Short turns are deliberate because they are easier for synthetic speech to render clearly. The authors connect this pattern to cognitive apprenticeship: modeling through the expert's reasoning, coaching through responses, and scaffolding through step-by-step support. They are careful not to claim a full implementation of cognitive apprenticeship. In particular, the support does not gradually fade as a learner becomes independent. The paper calls the system inspired by that framework, not equivalent to it. This distinction is good scientific hygiene. Borrowing a mechanism from a learning theory is not the same as validating the complete theory. It also points to a product question. A generic novice can ask tidy questions, but a real learner's confusion may be messier, more personal, or specific to one missing prerequisite. The study later found exactly that boundary. Fifty-seven point three percent agreed that the novice's questions helped them understand, but only 39.8 percent said the questions resembled questions they themselves had. The novice functioned better as prepared scaffolding than as a faithful stand-in for every learner. ## 5. What the classroom comparison actually did The participants were 245 first-year students at one prefectural high school in Niigata, Japan. They experienced three inquiry-based lessons during 2025. First, in May, they watched an instructor-narrated lesson about education and career paths. Second, in June, they watched a single-speaker TTS lesson about how generative AI works. Third, in July, they watched a dialogue-TTS lesson about data-driven decision making. That sequence creates the study's central limitation. Format changed, but so did topic, time, and familiarity with the broader course. A student might rate the July lesson more highly because of the dialogue, because the topic was easier to connect with, because this was the third encounter, or because several causes worked together. Questionnaires after the two synthetic lessons produced 229 valid responses for the single-speaker session and 206 for the dialogue session. Because responses were anonymous, the researchers could not link one student's first questionnaire to the same student's second questionnaire. They therefore treated the detailed comparison as approximately repeated cross-sectional, even though many respondents probably belonged to the same cohort. At the end of the dialogue questionnaire, students also retrospectively rated all three video formats. Up to 183 respondents supplied matched values for three core items: ease of comprehension, ease of concentration, and overall evaluation. The measures were created for this study. They were informed by the ARCS motivation model, cognitive load theory, and engagement theory, but the authors did not validate full scales with reliability coefficients or factor analysis. So a result should be described as a difference on a specific questionnaire item, not as proof that an entire theoretical construct changed. ## 6. Was synthetic speech worse than the instructor? For the retrospective three-format comparison, average ratings were close. Ease of comprehension was 3.55 out of five for the instructor, 3.51 for single TTS, and 3.52 for dialogue TTS. Concentration was 3.29, 3.24, and 3.27. Overall evaluation was 3.42, 3.33, and 3.35. The Friedman tests found no statistically significant format differences on those three items. But a non-significant difference is not automatically evidence that formats are equivalent. The researchers therefore added TOST, pronounced toast, an equivalence test. Using a margin of plus or minus half a point on the five-point scale, each synthetic format fell inside the equivalence range relative to the instructor, with p values below point zero zero zero one. That sounds reassuring, but it needs precise wording. The analysis supports equivalence of the retrospective ratings within the chosen half-point margin. It does not prove equivalent instruction. Students were remembering lessons delivered one and two months earlier, the content differed, and the half-point margin was a conventional ten percent of the scale rather than a demonstrated minimum educationally meaningful difference. So the practical signal is that these synthetic lessons did not receive dramatically worse core experience ratings than the instructor lesson. That lowers one barrier to experimentation with TTS. It does not establish equal learning, equal warmth, or equal accessibility for every student. ## 7. Where dialogue looked better, and worse The detailed comparison between single and dialogue TTS produced the episode's most interesting pattern. On the statement, "I think I understood the main content," 69.0 percent of the single-TTS group gave a rating of four or five, compared with 78.2 percent in the dialogue group. The corrected q value was point zero two five, with a small effect size. On, "I think I can explain the main content to a friend," the proportions were 31.0 percent and 43.7 percent. The corrected q value was point zero two one, again a small effect. Dialogue also scored higher on whether students tried to deepen their thinking: 52.0 percent versus 61.7 percent, with corrected q equal to point zero four eight. The researchers tested twenty items, so they controlled the false discovery rate instead of treating every uncorrected p value as independent proof. For the item saying the lesson was enjoyable, the uncorrected comparison favored dialogue, but it did not remain significant after that correction. A supplementary proportional-odds model controlling for prior knowledge did find a dialogue advantage on enjoyment, with an odds ratio of 1.65 and corrected q of point zero two five. The strongest preference result was simpler. Among 154 respondents who chose the most enjoyable of the three video formats, 66.9 percent chose dialogue. Among those answering which format they would like to experience again, 50.5 percent selected dialogue, compared with 27.7 percent for the instructor video and 22.3 percent for single TTS. Preference is valuable for sustained use, but it is still preference. Now the cost. Single-speaker TTS was rated more natural: 55.5 percent gave it a four or five, versus 35.4 percent for dialogue. The difference survived correction with q below point zero zero one and a small-to-medium effect size. Single TTS was also easier to hear. The paper proposes that generating audio slide by slide produced small variations in voice quality. One narrator can survive those variations as a slightly inconsistent person. With two characters, variation makes speaker identity harder to track. In other words, dialogue may improve cognitive scaffolding while adding audio-processing work. ## 8. Limits, confounds, and the missing learning test One imbalance initially seems to strengthen the dialogue result. Sixty-six percent of respondents in the dialogue group said they knew nothing about that lesson's topic beforehand, compared with 34.1 percent in the single-TTS group. The authors added an ordinal model controlling for prior knowledge, and key dialogue-favoring items remained significant. That analysis helps with one confound. It does not remove the others. Lesson topic, order, date, repeated exposure to the course, and which students completed each anonymous questionnaire still changed or could not be linked. A covariate model cannot reconstruct the randomized identical-content experiment that was never run. The missing outcome is objective learning. Students were not given validated pretests, post-tests, delayed retention tests, or transfer problems tied to the content. The study therefore cannot say dialogue increased knowledge. Phrases such as "self-assessed comprehension" and "perceived learning effect" are not fussy disclaimers. They name what was measured. There are other boundaries. This was one school and one age group. The questionnaires were made for the project and their construct validity was not established. Recall may have smoothed differences in the retrospective comparison. Response rates differed between the synthetic lessons. The manuscript is a preprint, so peer review may lead to corrections or reframing. None of that makes the experiment worthless. Practical exploratory studies are useful when they identify a pattern worth testing. The mistake would be to promote the pattern into a causal law before the decisive test. ## 9. What you can take away Five takeaways, each labeled by how settled the evidence is. One. Well supported here: students' retrospective ratings of comprehension, concentration, and overall experience for the synthetic lessons were close to the instructor lesson within the study's chosen equivalence margin. Two. Suggestive: expert-and-novice dialogue may improve self-assessed understanding, confidence, active thinking, and willingness to return compared with one synthetic narrator. Three. Well supported here: dialogue created an audio-quality trade-off; the single synthetic narrator was rated more natural and easier to hear. Four. Still unknown: whether dialogue caused more knowledge, retention, or transfer, because the study changed content and order and did not administer objective learning tests. Five. Well supported as a design principle: human review is part of the proposed system, not an optional cleanup step. The evidence does not support removing educators from factual and pedagogical judgment. Where is the field heading? The evidence dossier suggests a move from asking whether TTS is acceptable toward evaluating whole AI-assisted learning systems: dialogue structure, personalization, production workflow, content integrity, and voice quality together. That direction is suggestive. The decisive evidence would be a preregistered randomized crossover study in which the same material is narrated in each format, order is counterbalanced, learners are linked anonymously, scoring is blinded, and tests measure immediate understanding plus delayed retention. ## 10. Skeptical checklist and exact source card Before repeating a headline, ask six questions. Was the paper peer reviewed? Not yet. Were formats randomized? No. Was lesson content held constant? No. Were the same students linked across detailed questionnaires? No. Was actual learning tested? No. Was the positive signal meaningless? Also no: several corrected self-report differences and a strong enjoyment preference justify a better-controlled study. For a teacher, the practical experiment is small. If synthetic narration makes a lesson feasible, test both a clean single voice and a carefully written dialogue. Make the novice ask real diagnostic questions. Keep character identity acoustically stable. Have a subject expert review every factual and instructional decision. Then measure a learning outcome, not only whether the format felt engaging. The academic source is "A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential," by Gendo Kumoi, Fumie Watanabe, Tota Suko, Takashi Ishida, Yuko Kuma, Manabu Kobayashi, and Shigeichi Hirasawa. Version one was submitted to arXiv on July 14, 2026, as arXiv 2607.12235. The paper states that it is a preprint submitted for consideration to IEEE Access and has not yet been peer reviewed. It is available under the Creative Commons Attribution 4.0 license. The work was supported by the Japan Society for the Promotion of Science J-PEAKS program, grant JPJS00420240017. The authors disclose that LLMs were used in the educational-content system under study and also for manuscript proofreading and translation, with accuracy checked by the authors. The final verdict: this is a useful exploratory result, not a verdict on AI teaching. Two synthetic voices may make a lesson feel easier to understand and more enjoyable, but they can also sound less natural. What we know is how students rated three different sessions. What we still need to know is whether a well-reviewed dialogue, compared fairly on the same lesson, helps them learn and remember more.