As the overview article noted, besides his work on the vowels of old Japanese, Arisaka Hideyo built up in a book, Onin-ron (1940), the very question of what a “sound” is. Let’s check its entrance with a familiar example.
Watching the “n” (ん) inside your mouth
Say the following three words slowly aloud, and observe what your mouth is doing while you say the “n” (ん).
- sanma — through the “n,” your lips are closed.
- santa — through the “n,” the tip of your tongue touches around the back of your upper teeth.
- sankaku — through the “n,” the back of your tongue is pressed against the roof of your mouth.
We write it the same “ん” and feel it as the same “ん,” yet what happens inside the mouth differs in all three. In actual sounds, they come close to [m] (a nasal with closed lips), [n] (a nasal with the tongue tip), and [ŋ] (a nasal with the back of the tongue). The “ん” at the end of a word like hon (“book”) becomes yet another nasal, resonating at the back of the mouth.
Physically, each “ん” is a separate sound. And yet speakers of Japanese don’t try to tell them apart; they treat them all as one “ん.”
Phonetic sound and phoneme
Here we distinguish two terms.
- Phonetic sound — the physical sound that actually comes out of the mouth. Each one differs subtly.
- Phoneme — a kind of sound that speakers of a language feel to be “the same” and treat as one unit.
The “ん” of sanma, santa, and sankaku are separate as phonetic sounds, but one “ん” as a phoneme of Japanese. We might say a phoneme is a unit of sound that helps distinguish words within that language. In Japanese, a difference between [m] and [n] rarely changes the meaning of a word (it is determined automatically by the following consonant), so [m], [n], and [ŋ] all gather into a single “ん.”
The same thing shows up in comparison with other languages. The consonant of the Japanese ra row is a distinctive sound, different from both English l and r, but for a Japanese speaker the difference between l and r does not distinguish words. That’s why Japanese speakers tend to hear English “light” and “right” as alike. A phoneme is the boundary line, different for each language, between “same” and “different.”
What Arisaka Hideyo was looking at
Onin-ron (1940) is a book that rebuilds exactly this distinction between phonetic sound and phoneme, citing examples from Japanese and many other languages. For example, on the Tokyo “u” sound, Arisaka says that although it looks at a glance like a back, unrounded vowel, in fact the tongue sits nearer the center of the mouth — a central sound (p. 68). It is clearly different from a sound like the English “oo,” made with lips rounded and pushed forward. The Japanese “u” is a distinctive sound that is hard to fit directly onto the vowels of European languages, even as a phonetic sound.
Arisaka then pushes this relation of phonetic sound and phoneme one step further, into the question of what a phoneme is in the first place. The “coin” analogy at its core, and the view of a phonemic system as “an institution inherited from our ancestors,” we read in the original in the next two articles.
Summary
| Point | Detail |
|---|---|
| Phonetic sound | The physical sound that actually comes out of the mouth; each one differs subtly |
| Phoneme | A kind of sound that speakers of a language feel to be “the same” and bundle as one |
| The “n” (ん) example | The “n” of sanma, santa, and sankaku are separate phonetic sounds but one phoneme of Japanese |
| Differences between languages | What counts as “the same” varies by language (the Japanese ra row versus English l and r) |
In a nutshell — the “same sound” we feel by ear is not physical sameness but sameness in the sense of lying within a boundary line drawn by the language called Japanese. That boundary line is the phoneme, and Arisaka asked what it is.