Listening Beyond the Words

– Alban Voppel –

Speech can reveal a lot about how someone is doing, and for people living with schizophrenia, this is especially true. In a recent study at the Centre of Excellence in Youth Mental Health (CEYMH), we used artificial intelligence (AI) to analyse short pieces of speech from people with and without schizophrenia-spectrum disorders, and discovered that the AI did not need to look at what people said, but simply at how their voice sounded.

Imagine asking someone, “How are you doing?” They answer, “Fine.” The word itself tells you very little. But the voice might tell you more. Was the answer flat? Was there a long pause before it? Did the person’s voice rise and fall naturally, or did it stay almost the same?

In everyday life, we pick up on these signals all the time and we can tell whether someone is sad or happy. However, could a computer learn to recognize these same patterns and help with mental health diagnosis?

alban voppel AI mcgill mental health psychosis voice
This image explains how voice recordings are transformed into data that can be analyzed by AI

In 2025, together with Gleb Melshin and other researchers from the CEYMH, we analyzed voices from more than 300 people with schizophrenia-spectrum disorders as well as healthy volunteers, and found that our AI model correctly identified those with the disorder in 88% of cases, and reached the same accuracy when detecting more severe symptoms.

These disorders can come with symptoms that are hard to see, even for those who experience them. Take blunted affect, for example. It is one of the more severe symptoms of the illness, and it refers to a reduction in emotional expression. Yet a person may not themselves realize having a flatter voice or speaking less, while family members may notice that “something has changed,” without being able to know what. Clinicians can observe these signs in an interview, but it can be hard to distinguish from depression or having symptoms in general. 

How did the AI analyse the voice?

We divided the recordings into 10-second clips, long enough for a sentence, and transformed each clip into an image called a spectrogram. This technique is like a fingerprint of speech, it shows how the sounds of a person’s voice change over time. Different speech patterns create different images, allowing computers to look for subtle features that may be difficult to notice.

Some AI models are very good at analyzing images. Instead of learning to recognize objects such as cars or cats, the computer model was trained to recognize hidden patterns in human speech. 

However, can we be sure that the computer is not making a mistake?

Making sure the AI was listening to the right things

AI can sometimes learn the wrong information. For example, it might accidentally focus on background noise, recording quality, or other detail that has nothing to do with the voice of a person. In this study, we ran additional checks to understand what patterns the model learned. It seems that the model focused on parts of the spectrogram that matched human speech frequencies, rather than irrelevant audio features.

This suggests that the model actually learned information from human speech, including patterns that reflect blunted affect and information about diagnosis.

One surprising detail is that the model was not simply finding general illness, it seemed to look for different clues in the voice depending on the symptoms studied. This suggests that the voice carries different kinds of clues, depending on what we are trying to understand, and a computer model can detect these.

There are still important limitations. All speech in this study was English, so the same approach needs to be tested in other languages to see if it works in the same way. Moreover, medication, illness stage, and recording conditions may also affect the voice, and the model needs to be tested in new clinics and real-world settings before we know how useful it will be in practice.

Tomorrow, your voice could become a clinical tool

These findings do not mean that a computer can diagnose schizophrenia from someone’s voice, nor should they be used that way. Mental health diagnosis is complex and depends on a person’s experiences, behaviour, functioning, and clinical context. A voice recording is only one piece of information, but it could become a useful extra tool. Think of it as a thermometer; it does not replace a doctor, and does not explain why someone is sick, but it gives a clear measurement that can help detect illness and can track changes. In the future, speech analysis might help clinicians track changes over time, notice when symptoms are getting worse, or measure symptoms that are hard to describe.

With this new research, computers may help us listen in a more detailed and consistent way someone speaks. Not to replace human care, but to support it, especially for symptoms that are easy to miss or hard to detect. Through careful listening, supported by responsible technology, this research may one day contribute to more personalized care, earlier intervention and better outcomes for people living with psychosis.

Because in the end, the answer to ‘How are you doing?’ might carry more information than we ever imagined.

Discover the original scientific article

Taking a look at your speech: identifying diagnostic status and negative symptoms of psychosis using convolutional neural networks

NPP-Digital Psychiatry and Neurosciences, 2025

Discover the author

Alban Voppel research coordinator, post doctorat student from Mcgill university at the center of excellence in youth mental health

Dr. Alban Voppel is an Assistant Professor at McGill University and research coordinator on the MOTS+ project at the CEYMH, within the Douglas Research Centre. 

His work focuses on how speech, language, and artificial intelligence can help us better understand psychosis and related mental health symptoms.

Glossary

Schizophrenia-spectrum disorders: group of mental health conditions that can involve hallucinations, disorganized thoughts, or changes in motivation and expression

Spectrogram: image that represents someone’s voice, showing how it changes over time (higher or lower pitch, louder or softer)

Blunted affect: reduction in emotional expression where a person speaks in a flat, monotone voice and shows fewer facial expressions than usual, even in situations that would normally trigger an emotional reaction