Research
Chat-based language learning, by the numbers
We analyzed millions of words of real chat, conversation, and written text in English, Japanese, and Spanish, and modeled what an hour of chat practice actually gives you. Here's what holds up, and what doesn't.

We make a strong claim about chat at MiraChat: it's slow enough to read and write, fast enough to stay engaging, full of the vocabulary you actually need, and it leaves a transcript you can study. Claims like that are easy to make and hard to check, so we tried to check them.
We pulled public research on reading, typing, and turn-taking, analyzed real chat corpora next to spoken conversation and written text in three languages, and built a simple model of what an hour of practice looks like in different formats. Some of the results strongly support chat. A couple don't, and we've included those too.
The short version
1. Chat gives you time to think: about 100 times more
In spoken conversation, the gap between one person finishing and the next person starting averages about 200 milliseconds across languages.1 Producing a spoken reply takes longer than that, over 600 milliseconds just to plan a short utterance, so people plan their answer while the other person is still talking.2 That's hard enough for native speakers. For learners it's often impossible.
To see how much time text chat gives you, we looked at RealPersonaChat, a corpus of about 13,600 Japanese text chats between real people, with a timestamp on every message.3 Across about 400,000 replies:
That's roughly 100 times the time you get in speech, and it's what native speakers take. Nobody in these chats expected an instant answer, which means a learner can take the time to reread a message, look up a word, or fix a verb ending, and the conversation still feels normal.
2. Chat runs on everyday vocabulary
The second claim is that chat uses practical, high-frequency language. To test it, we took samples of up to 300,000 words from different kinds of text and asked a simple question: if you knew the N most common words in the language, what share of the text would you recognize?
We compared one-on-one text chat, spoken conversation, film and TV subtitles, fiction, news, academic writing, and Wikipedia, using standard frequency lists for each language.4
Two things stand out.
Text chat has the vocabulary profile of spoken conversation. In both English and Japanese, one-on-one chat lands right next to recorded speech, and far from written text. Learning from chat means learning the vocabulary of conversation.
The gap matters more than it looks. A few percentage points of coverage sounds small, but it's the difference between a comfortable read and a frustrating one. Here's the same data framed as unknown words, assuming you know the 3,000 most common words:
| Text type | Unknown words per 100, English | Unknown words per 100, Japanese |
|---|---|---|
| One-on-one text chat | ~9 | ~11 |
| Spoken conversation | ~11 | ~11 |
| Wikipedia | ~21 | ~26 |
At around nine unknown words per hundred, you can follow a chat with a few lookups. At more than twenty, you're decoding rather than reading.
For context, research on reading suggests learners need to know about 95–98% of the words in a text to understand it comfortably,5,6 and estimates put the vocabulary needed for spoken English at around 6,000–7,000 word families, versus 8,000–9,000 for written text.7 (Our counts are word forms, not word families, so they aren't directly comparable with those thresholds, but the direction is the same.)
One important caveat
"Online text" is not automatically easy. Public chat rooms from the 2000s (the NPS Chat corpus) and Spanish reply tweets, both full of slang, usernames, creative spelling, and topic hopping, actually needed more vocabulary than Wikipedia to reach high coverage. The easy profile belongs specifically to one-on-one conversational chat: two people talking about their lives. That's the kind of chat worth learning from.
3. Chat is shaped like conversation
Beyond vocabulary, chat has a distinctive shape, and almost every feature of that shape helps learners:
| Measure | One-on-one text chat | News, academic, Wikipedia |
|---|---|---|
| Words per sentence (English) | 6–9 | 15–20 |
| Messages or sentences with a question | 19–30% | about 1% |
| I and you pronouns per 1,000 words | 120–155 | 3–8 |
| Distinct words per 10,000 (English / Japanese) | ~1,600 / ~1,800–2,100 | ~3,000 / ~3,200 |
Short sentences are easier to parse and to imitate. Questions demand an answer, which is what makes chat interactive. Constant I and you means the language is about the people talking, so the language you read is language you can turn around and use about yourself. A smaller set of distinct words means the same words keep coming back.
There's one more nuance for Japanese learners. Chat between strangers in the RealPersonaChat corpus is heavily polite, with about 110 です/ます forms per 1,000 words versus about 10 in recorded conversation between friends. Chat teaches whichever register the relationship calls for. If you want casual Japanese, you need partners who talk to you like a friend.
4. What an hour of practice gives you
Next we built a simple model comparing an hour of one-on-one text chat with an AI partner, a one-on-one spoken tutor lesson, a group class, and a gamified app session. It combines published anchors (reading speed,8 phone typing speed,9 the 200 ms turn gap, and a video study of 105 classrooms in which teachers talked for 51% of lesson time and all students combined for 23%10) with clearly labeled assumptions about message length, lookup time, and so on. These are estimates, not measurements. Here's the middle scenario:
| One hour of… | Words you read | Words you hear | Words you produce | Your turns | Time to plan each turn | Transcript? |
|---|---|---|---|---|---|---|
| Text chat with an AI partner | ~1,000 | 0 | ~400 | ~40 | ~60 s | Yes |
| Spoken tutor lesson | 0 | ~3,500 | ~2,500 | ~320 | ~0.2 s | No |
| Group class | ~500 | ~4,800 | ~430 | ~54 | ~0.2 s | No |
| Gamified app | ~1,700 | ~960 | ~240 (mostly given prompts) | 0 | n/a | Partial |
Note what this table does not say: chat doesn't win on raw volume. A good spoken tutor gets far more words out of you per hour. We think that's worth being upfront about.
What chat wins on is a different set of things:
- Planning time. Each of your chat turns comes with about a minute to read, look things up, and compose. That's when you can actually notice grammar and try new structures, instead of falling back on what you already know.
- Self-generated output. Every word you write in a chat is your own sentence about your own life, unlike an app, where most typed answers translate a prompt someone else wrote.
- Turns versus a class. In the middle scenario, a student in a group class speaks for about 4.5 minutes per hour. An hour of chat gives you about 40 turns of your own.
- A permanent record. Every lookup, correction, and new phrase is still there afterward to review.
- Availability. A tutor costs money and needs scheduling. A class runs on someone else's timetable. Chat fits into ten spare minutes, as often as you like.
5. What a year of chat adds up to
Finally, we modeled 30 minutes of chat a day for a year (about 182 hours) by simulating the words a learner would read and write, drawn from the frequency patterns of real conversational text. In the middle scenario:
Ten encounters isn't a magic threshold. Research finds that the number of times you meet a word is moderately linked to learning it (an average correlation of about .34 across 26 studies), and that one or two encounters rarely do the job.11,12 But a year of daily chat gives you repeated, in-context exposure to the core of the language, without a single word list.
For scale, 182 hours is about a quarter of the classroom time the US Foreign Service Institute estimates for Spanish (600–750 hours) and about 8% of its estimate for Japanese (about 2,200 hours).13 Chat alone won't get you there, but it's a serious amount of practice to fit into otherwise idle half-hours.
6. What the research says about chat versus other practice
Classroom studies have compared text chat with face-to-face practice for about 30 years. The most useful findings:
- More output. In one early study, students produced two to four times as many sentences, and twice as many turns, in chat discussions as in oral discussions.14 Participation was also more evenly spread, with quieter students contributing more.15
- Transfer to speaking. In one small quasi-experimental study, students who spent half their class time in chat sessions scored higher on oral proficiency than students taught entirely face to face.16
- Less anxiety. In a study comparing text and voice chat, speaking anxiety fell only in the text-chat group.17
- Overall, about as effective as face-to-face. Meta-analyses find chat-based interaction roughly as effective as face-to-face interaction, with some advantage for writing and production.18,19
That last finding is the honest summary: text chat isn't a magic shortcut. It's about as effective as practicing in person, and much easier to get a lot of. For most learners the constraint isn't which kind of practice is best. It's how much practice they can actually fit in.
What we're not claiming
- That chat beats speaking. If your goal is speaking, you'll need to speak eventually. Chat builds the vocabulary and structures you'll speak with.
- That chat gives more practice per hour than a tutor. On raw word count, it doesn't.
- That all online text is easy. Only one-on-one conversational chat shows the easy, everyday profile.
- That our model numbers are measurements. They're estimates built on stated assumptions, and we'll replace them with real usage data as we collect it.
How we did this
Corpora: PersonaChat, EmpatheticDialogues, the NPS Chat corpus, a Switchboard sample, the Brown corpus, English Wikipedia and OpenSubtitles (English); RealPersonaChat, JPersonaChat, JEmpatheticDialogues, the Nagoya University Conversation Corpus (NUCC), Japanese Wikipedia and OpenSubtitles (Japanese); and reply tweets, OpenSubtitles and Wikipedia (Spanish; we couldn't find a free one-on-one Spanish chat corpus). We lowercased word forms; removed punctuation, numbers, emoji, URLs, and usernames; excluded proper nouns; tokenized Japanese with MeCab; and ranked words by the wordfreq frequency lists. Results were checked against subtitle-only and Wikipedia-only frequency lists. The English chat advantage holds under both, though it narrows under a Wikipedia-only list. Timing figures come from RealPersonaChat timestamps, excluding gaps over 10 minutes. The practice and one-year figures come from a model with low, middle, and high scenarios. We report the middle one here.
Sources
- Stivers, T., et al. (2009). Universals and cultural variation in turn-taking in conversation. PNAS, 106(26), 10587–10592. https://doi.org/10.1073/pnas.0903616106 ↩
- Levinson, S. C., & Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6, 731. https://doi.org/10.3389/fpsyg.2015.00731 ↩
- Yamashita, S., et al. (2023). RealPersonaChat: A realistic persona chat corpus with interlocutors' own personalities. PACLIC 37. Corpus: https://github.com/nu-dialogue/real-persona-chat ↩
- Speer, R. wordfreq: word frequencies in many languages. https://github.com/rspeer/wordfreq ↩
- Hu, M., & Nation, I. S. P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language, 13(1), 403–430. https://eric.ed.gov/?id=EJ626518 ↩
- Laufer, B., & Ravenhorst-Kalovski, G. C. (2010). Lexical threshold revisited: Lexical text coverage, learners' vocabulary size and reading comprehension. Reading in a Foreign Language, 22(1), 15–30. https://eric.ed.gov/?id=EJ887873 ↩
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82. https://doi.org/10.3138/cmlr.63.1.59 ↩
- Brysbaert, M. (2019). How many words do we read per minute? A review and meta-analysis of reading rate. Journal of Memory and Language, 109, 104047. https://doi.org/10.1016/j.jml.2019.104047 ↩
- Palin, K., Feit, A. M., Kim, S., Kristensson, P. O., & Oulasvirta, A. (2019). How do people type on mobile devices? Observations from a study with 37,000 volunteers. MobileHCI '19. https://doi.org/10.1145/3338286.3340120 ↩
- Helmke, T., Helmke, A., Schrader, F.-W., et al. (2008). Die Videostudie des Englischunterrichts. In DESI-Konsortium (Ed.), Unterricht und Kompetenzerwerb in Deutsch und Englisch. Beltz. https://www.pedocs.de/volltexte/2013/3521/pdf/Helmke_Helmke_Schrader_Videostudie_2008.pdf ↩
- Uchihara, T., Webb, S., & Yanagisawa, A. (2019). The effects of repetition on incidental vocabulary learning: A meta-analysis of correlational studies. Language Learning, 69(3), 559–599. https://doi.org/10.1111/lang.12343 ↩
- Webb, S. (2007). The effects of repetition on vocabulary knowledge. Applied Linguistics, 28(1), 46–65. https://doi.org/10.1093/applin/aml048 ↩
- Foreign Service Institute, U.S. Department of State. Foreign language training: language learning difficulty for English speakers. https://www.state.gov/foreign-language-training ↩
- Kern, R. G. (1995). Restructuring classroom interaction with networked computers: Effects on quantity and characteristics of language production. The Modern Language Journal, 79(4), 457–476. https://doi.org/10.1111/j.1540-4781.1995.tb05445.x ↩
- Warschauer, M. (1996). Comparing face-to-face and electronic discussion in the second language classroom. CALICO Journal, 13(2–3), 7–26. https://doi.org/10.1558/cj.v13i2-3.7-26 ↩
- Payne, J. S., & Whitney, P. J. (2002). Developing L2 oral proficiency through synchronous CMC: Output, working memory, and interlanguage development. CALICO Journal, 20(1), 7–32. https://doi.org/10.1558/cj.v20i1.7-32 ↩
- Satar, H. M., & Özdener, N. (2008). The effects of synchronous CMC on speaking proficiency and anxiety: Text versus voice chat. The Modern Language Journal, 92(4), 595–613. https://doi.org/10.1111/j.1540-4781.2008.00789.x ↩
- Ziegler, N. (2016). Synchronous computer-mediated communication and interaction: A meta-analysis. Studies in Second Language Acquisition, 38(3), 553–586. https://doi.org/10.1017/S027226311500025X ↩
- Lin, W.-C., Huang, H.-T., & Liou, H.-C. (2013). The effects of text-based SCMC on SLA: A meta-analysis. Language Learning & Technology, 17(2), 123–142. ↩