Prosody-Driven Semantic Interpretation in LLMs
Problem
In speech, the same words can be a statement or a question depending only on pitch. A text-trained language model does not receive that cue unless it is supplied as an auxiliary signal, which means a model that looks strong on transcripts can still miss the speech act that a listener hears immediately. The hard case is the one where the words cannot help: the lexical content is identical, and the only remaining information is the pitch contour. The problem is that current LLMs may fall back to chance once the text is no longer diagnostic.
Method
We recorded 1,200 utterances from 12 speakers, 100 files each, split evenly into 600 declarative and 600 interrogative productions of the same sentences, so that the words cannot decide the label. Pitch contours were extracted in Praat as time-aligned series. Two pipelines were compared. Alpha transcribes the audio with Whisper, adds summary prosodic statistics from Librosa, and asks a model to label a broad tone; that pipeline was limited by transcription errors and did not stay speech-native. Beta, the reported comparison, sends the pitch contour itself to the model as a time-pitch series, with no transcript, and asks the model to decide interrogative versus declarative from the rising or falling pattern. The compared models are Llama 3 8B and Qwen 2 7B, run locally through Ollama, and Gemini 1.5 through an API. Accuracy, precision, recall, and F1 are computed on the binary speech-act label, with chance at 50% on this balanced set.
Results
Gemini 1.5 reaches 70.8% accuracy, with precision 0.728, recall 0.665, and F1 0.695, and a speaker-level mean of 0.708 ± 0.111. Qwen 2 7B reaches 51.8%. Llama 3 8B reaches 50.7%. The two local models are therefore essentially at chance when the only cue is pitch. The errors are not symmetric. Llama 3 predicts declarative far too often, sending 435 of 600 interrogative utterances to the declarative class. Qwen 2 does the reverse, sending 407 of 600 declarative utterances to the interrogative class. Gemini is well above chance and still far from reliable speech-act classification, which is consistent with a broader claim: current text-centric models do not automatically inherit the prosodic distinctions that ordinary spoken language depends on.