Dissertation Cross-language voice variation and talker identity What happens to a voice when speakers switch languages.

My dissertation project asks what happens to a voice when a speaker switches languages and speech styles. Speakers were recorded in Korean and English, in read and spontaneous styles. One part of the study examined listeners’ ability to judge whether two clips came from the same person; the other focused on analyzing the voices directly.

The perception study tested 276 listeners from three language backgrounds, asking them to judge whether pairs of voices were from the same talker. Pairs differed in language and style.

The talker effect across languages and speech stylesMean same-different ratings with 95% confidence intervals. Hollow points are pairs that were the same talker, filled points are pairs that were two different talkers.
Study 1 · Perception. Manuscript in preparation. How different the talker sounded, with 95% confidence intervals; rows are ordered by the size of the talker effect. The talker effect was greatest when the speech style was held constant, and it was larger for female talkers. Language made a different contribution: pairs that crossed languages received higher “different talker” ratings, while same‑language pairs made people sound more similar.

The second study examined nine acoustic measures of voice from 28 Korean heritage speakers in Korean and English, across speech styles.

Where the language effect appears across nine voice measuresLanguage effects concentrate in read speech, and women mark the contrast in pitch while men mark it in resonance and phonation.
Study 2 · Three dimensions. Manuscript in preparation. Nine acoustic measures capture pitch, resonance (higher formants), and phonation, and the cells show where Korean and English diverge. Two patterns emerge: the contrast appears almost entirely in read speech, and women carry it in pitch while men carry it in resonance and phonation.

A third analysis reduced 26 acoustic measures of voice to eight voice components and then tested which factors move a voice within that space.

How much each factor moves the voice, component by componentPartial eta squared for every effect across the eight retained principal components of voice. Gender is the darkest row, then language.
Study 3 · Voice structure. Manuscript in preparation. Partial η² across the eight components shows talker gender as the strongest effect, language as the next, and speech style as significant chiefly in interaction terms.
Sound systems Multilingualism How multilingualism shapes the perception and production of speech.

One of my main research interests concerns how multilingualism influences voice quality, articulatory gestures, speech perception and production (including both categorical and within-category dimensions), sound change, and language acquisition. I am particularly interested in how multilingual speakers navigate and manage distinct phonetic and phonological systems across languages.

In Cho, Jongman, Wang & Sereno (2020), Multi-modal cross-linguistic perception of fricatives in clear speech, published in the Journal of the Acoustical Society of America, I looked at how listeners from three language backgrounds identify English fricatives when the signal arrives as audio alone, as video alone, or as both together — and whether speaking clearly helps.

Where clear speech helps fricative identificationEffect of clear speech on fricative identification by modality, sibilance and listener language background. Clear speech helps most in audio-only sibilants and reduces audio-visual sibilant accuracy for English and Korean listeners.
Clear speech does not confer a uniform benefit. From Cho, Jongman, Wang & Sereno (2020). Clear speech was most reliably helpful for audio-only sibilants, in all three listener groups; in visual-only modality, it helped only the two non-English groups. For audio‑visual perception of sibilants, it actually impeded accurate perception, lowering accuracy for both English L1 and Korean L1 listeners.

In Cho (2023), What does it mean to sound Korean-Canadian, eh? A comparative study on Canadian English vowel space, presented at the Acoustical Society of America, I looked at whether the vowel space of Korean heritage speakers shifts when they switch between their two languages.

Vowel space of Korean heritage speakers in English and KoreanLobanov-normalized F1 by F2 vowel spaces, men in one panel and women in the other. The two languages largely overlap. The women back their /u/ in Korean; for the men the two /u/ tokens sit almost on top of each other.
The vowel spaces overlap, though women have a backer Korean space. Most vowels align across languages. Both groups back /o/ in Korean; only women back /u/, giving their Korean a more distinct “Korean” sound. Men’s /u/ tokens are almost identical.

In Cho & Munro (2017), F0, long-term formants and LTAS in Korean-English bilinguals, presented at the Phonetic Society of Japan, I looked at how Korean-English speakers differ in pitch, formants, and energy spread across the spectrum.

Korean is quieter than English at high frequenciesAbove 2 kHz the same bilingual speakers produce Korean with lower intensity than English, by 3 to 5 dB between 2 and 4 kHz and 5 to 10 dB higher up. Six of the ten speakers showed the pattern.
Korean has reduced high-frequency energy. The languages align below 2 kHz, but Korean shows a drop-off above that point, with the gap widening as frequency increases. The reduced high-frequency energy likely reflects Korean’s sparse fricative system or its articulatory setting.
Sensorimotor systems Speech and hearing disorders How neurodegenerative conditions reshape articulation and voice.

I am interested in how speech disorders — for example, neurodegenerative conditions such as Parkinson’s disease — affect articulation and voice quality. I am also interested in the relationship between speech and other motor domains, such as fine and gross motor control, and how these systems interact in both healthy and clinical populations.

In Diep, Cho, Shamei, Liu & Gick (2024), we used mPower data (Bot et al., 2016) to compare voice, tapping, and walking data in participants with and without Parkinson’s.

What predicts a Parkinson’s diagnosis across three motor tasksIn the voice, pitch rises with Parkinson’s and harmonics-to-noise ratio falls. In finger tapping and walking every predictor falls.
Speech and the rest of the motor system. From Diep, Cho, Shamei, Liu & Gick (2024). A logistic regression predicting a Parkinson’s diagnosis from features taken at the onset of three tasks in the mPower mobile dataset.
MDS cluster centroids for the three motor tasksAll three tasks separate patients from controls. Only the voice also separates the onset of a movement from its middle in both groups; in finger tapping that gap survives only in the participants without a diagnosis, and in walking it is absent.
The clusters, in MDS space. From Diep, Cho, Shamei, Liu & Gick (2024), Figure 1. All three tasks separate PD patients apart from controls. However, only voices further separate the start of a movement from its middle.
Theory into practice Applied phonetics Turning phonetic research into teaching, outreach, and assessment.

I aim to bridge the gap between linguistic theory and real-world practices by exploring how insights from phonetic research can be translated into effective tools for pronunciation instruction, forensic linguistics, and speech-based AI technologies.

That also meant introducing speech research to people who had never encountered the field before. From 2023 to 2025, I gave a talk called Linguistics in the World: A Uniquely Human Ability through SFU’s FASS in the Classroom programme, visiting high school classrooms across Greater Vancouver. The talk makes the case that linguistics is a science that studies the language around us, uses scientific methods and evidence, and provides insights that shape technology, illuminate social phenomena, and support real‑life jobs.

Three demos from the Linguistics in the World talkA spiky shape and a round one for bouba and kiki; a knit hat that goes by toque, beanie or tuque depending where you are; and a spectrogram of a spoken word.
Linguistics in the World: A Uniquely Human Ability. SFU FASS in the Classroom, 2023–2025. Three of the demos. Every stop starts from something the room already knows and ends somewhere they did not expect to be.

I also worked directly with classrooms and contributed to teaching practice by assessing aspiring teachers in SFU’s LING 363, the course leading to the TESL Canada credential.