Type “Nahuatl” into a technology pitch and a language can become a checkbox. In practice, speakers use distinct varieties with their own communities, pronunciations and linguistic histories. Mozilla’s Common Voice releases make that diversity visible in an unusually plain place: the dataset catalogue. The 2026 scripted-speech edition includes separate entries for Central Puebla, Western Sierra Puebla and Orizaba Nahuatl. A single label would erase the difference before a model ever heard a voice.
The Central Puebla Nahuatl datasheet identifies the variety with the locale code ncx and describes read-speech recordings meant for speech-recognition research. A separate Orizaba Nahuatl datasheet uses nlv. The catalogue also lists Western Sierra Puebla Nahuatl under nhi. These are not three brand names for an interchangeable file. They are labels that tell a researcher which speech community a dataset purports to represent and where a model’s results should be evaluated. The files are offered under CC0 terms, but a legal license does not settle every question about responsible use or community benefit.
Speech recognition starts with recordings and text aligned closely enough for a system to learn the relationship between sound and words. That input is expensive in human time. The Central Puebla datasheet reports 9,509 clips from 41 speakers, while the Orizaba datasheet reports 6,844 clips from 16 speakers. Those are corpus counts, not a measure of how well a deployed system understands anyone. Choosing text, contributing speech and validating recordings all require human work. A dataset page can count bytes, but it cannot make that work interchangeable across language varieties. Nor does the existence of a dataset mean there is a reliable, production-ready assistant for every speaker of that variety. The distance between a corpus and a dependable service can be large.
The detail matters beyond linguistics. When a company says its product “supports Nahuatl,” what exactly has it tested? Which variety? Against whose speech? In what acoustic setting, with what error rate? Does the interface allow people to reject a wrong transcription, and is the correction useful to their community? Without those answers, a broad language label becomes a marketing shortcut. Precision in naming is a first protection against that shortcut, even though naming alone cannot guarantee performance.
Common Voice’s catalogue is a starting point, not a completed solution. It offers material that researchers can inspect and build upon, while its separate entries make a better question possible. The goal is not to force the diversity of Nahuatl speech into one marketable token. It is to build tools whose makers can say plainly which voices they heard, which they missed, and who had a say in the difference.