At the end of part 1 of the walkthrough we touched Sapir’s “sound pattern.” Here it is up close — chapter 3, and the road to phonology.

1. Objective sounds, and the sound in the mind

what the speaker feels: all the same "t", "s" the t of teem the t of sting the s of heads the actual sounds: teem = a t with a full breath-release; sting = a t whose release is inhibited by the s the s of heads is really a voiced [z]; the ea of meat is shorter than the ea of mead
The speaker feels "the same sound," but objectively several distinct sounds are in use. Diagram: L/LAB

Speakers feel their language is built of “a comparatively small number of distinct sounds.” Under analysis, the sounds actually used are far more numerous. “Probably not one English speaker out of a hundred has the remotest idea that the t of a word like sting is not at all the same sound as the t of teem” — the first has its breath-release inhibited by the preceding s. In the same way the final s of heads is a voiced [z], and the ea of meat is shorter than the ea of mead.

Yet the speaker treats these as one sound. This is Sapir’s point:

Back of the purely objective system of sounds … there is a more restricted “inner” or “ideal” system … a finished pattern, a psychological mechanism. (ch. 3)

2. The pattern changes more slowly than the sounds

older language A ← the sound content is replaced generation by generation → present language A′ (descendant) not one shared sound. And yet — the number, relation and function of the points (the pattern) can be the same. The pattern's rate of change is "infinitely less rapid" than that of the sounds.
The sound content can be wholly replaced while the skeleton of the pattern stays. So closely related languages can share the same pattern with no sound in common. Diagram: L/LAB

The ideal sound pattern “may persist as a pattern — involving number, relation, and functioning of phonetic elements — long after its phonetic content is changed.”

Two historically related languages or dialects may not have a sound in common, but their ideal sound-systems may be identical patterns. (ch. 3)

Its rate of change is “infinitely less rapid than that of the sounds as such.” So a language is characterised “as much by its ideal system of sounds … as by a definite grammatical structure.”

3. You can only write down the “points in the pattern”

Sapir taught Native speakers to write their own languages. The result was always the same.

actual speech (a continuous rumble) the pattern in the mind (a grid of points) writing it down = transcribing an "ideal flow of phonetic elements"
The speaker hears the rumble only through the pattern in the mind. Transcription becomes a copy of an "ideal flow." Diagram: L/LAB

Sapir’s interpreter seemed to be “transcribing an ideal flow of phonetic elements which he heard … as the intention of the actual rumble of speech.”

4. This became the phoneme

What Sapir called the “ideal system of sounds” crystallised, in later linguistics, into a unit: the phoneme.

Baudouin de Courtenaythe "sound in the mind" — the phoneme idea (1870s–) Sapir (US) the Prague School Trubetzkoy / Jakobson phonology(distinct from phonetics) A phoneme is fixed not by its own content but by its oppositions in the system. e.g. English /p/ is placed by "voiced or voiceless / labial or alveolar / nasal or not."
The "sound in the mind" goes back to Baudouin de Courtenay; through Sapir and the Prague School (Trubetzkoy, Jakobson) it became "phonology," which treats sounds as a system of oppositions. Diagram: L/LAB

The t of sting and the t of teem fall on the same point in English (they never distinguish meaning — they are allophones). In a language where aspiration does distinguish meaning (Hindi, Thai), they are different points. A “sound” is a physical fact; which sounds count as the same is decided by each language’s pattern.

5. The same intuition as Saussure

That a phoneme is fixed “not by its content but by its oppositions” is exactly Saussure’s idea that the value of a sign is fixed by difference. Saussure (1916) and Sapir (1921), independently, reached the same intuition: a unit of language is not a substance but a position in a system. The Prague School gave that intuition its most precise form, in the domain of sound.

In sum

Objective sounds are more numerous than speakers thinkthe t of sting ≠ the t of teem; the s of heads is [z]
Behind them is an “ideal sound pattern”unconscious, yet more accessible to consciousness than the raw sounds
The pattern changes far more slowly than the soundsrelated languages can share a pattern with no sound in common
Speakers hear sound only through the points of the patterntranscription = a copy of an “ideal flow”
This idea became the phonemephonetics (the sounds) split from phonology (the system of oppositions); a phoneme is a point on a grid

The last piece of the Sapir feature is language and literature (chapter 11) — what translates and what does not, and prosody as an outcrop of a language’s dynamics. For the overview, see “Reading Edward Sapir’s Language”.