Sign in

Auditory-Visual Speech Association (AVISA)

@avsp.bsky.social
1.2K followers 989 following 764 posts

The official(ish) account of the Auditory-VIsual Speech Association (AVISA) AV 👄 👓 speech references, but mostly what interests me avisa.loria.fr

PostsRepliesMedia
Reposted by Auditory-Visual Speech Association (AVISA)
Josh McDermott @joshhmcdermott.bsky.social · 8h
Postdoc opening in my lab to work on NeuroAI approaches to improve prosthetic devices for human hearing. We offer a great training environment and strong mentorship. Prior expertise in machine learning and signal processing are essential. Apply here: careers.peopleclick.com/careerscp/cl...
careers.peopleclick.com
Postdoctoral Associate
MIT - Postdoctoral Associate - Cambridge MA 02139
0139
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 21h
Rethinking musicians as a model for plasticity and transfer www.cell.com/trends/cogni... Proposes repositioning correlational research from inferring training effects to constraining causal accounts, revealing neurocognitive organization & guiding studies with genetically informed designs 🎵 🧠 🤹‍♀️
cell.com
Rethinking musicians as a model for plasticity and transfer
Differences between musicians and nonmusicians are routinely attributed to training-induced plasticity and transfer, even though pre-existing factors provide equally plausible explanations. We propose...
060
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 06/10/2026
And this one
AVSP 1998 4-6 December 1998 Terrigal, Sydney, Australia
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 05/10/2026
AVSP 2026 has gone, so let's go back to where it started
AVSP 1997 
26-27 September, 1997
Rhodes, Greece
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 05/10/2026
Highlights
Speech comprehension requires the brain to transform a transient acoustic input into stable linguistic representations across multiple levels of abstraction.
The present article introduces three interdependent coding properties that meet this challenge: persistent, parallel, and time-stamped encoding.
Persistence enables prolonged access to the speech input.
Parallel processing supports simultaneous encoding of a recent history of inputs, across multiple levels of representation, using a dynamic code.
Time-stamping preserves the relative temporal order of inputs to support further structure building.
Together, these properties provide a neural coding architecture that supports robust construction of linguistic structure during speech comprehension.
Abstract
The acoustic signal of speech is fleeting and unfolds linearly; yet listeners must derive temporally extended and hierarchically organized linguistic structures to comprehend it. This article reviews evidence that the human brain transforms the transient auditory signal into neural representations that are persistent, parallel, and time-stamped. First, neural persistence allows information to remain available after it has disappeared from the acoustic input. Second, parallel encoding of recent inputs across multiple levels of language structure enables composition and interactions across the speech hierarchy. Third, dynamic neural representations encode both content and elapsed time, providing a flexible mechanism for preserving sequence order. Together, these representational properties reveal how the human brain transforms a continuously disappearing signal into temporally extended linguistic structures.
010
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 05/10/2026
Time after time: neural codes of speech comprehension pubmed.ncbi.nlm.nih.gov/42830249
152
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 24/09/2026
📚 Citation Classic "Word association norms mutual information and lexicography" Church & Hanks (1990) Citations: 7600+ Popularized mutual information for finding... 🔗 scholar.google.com/scholar?q=Church… #SpeechScience
022
Reposted by Auditory-Visual Speech Association (AVISA)
wimpouw @wimpouw.bsky.social · 21/09/2026
Here we go! We can finally share what Sharjeel Shaikh @sharjeelshaikh has been leading the work on with us! Consider participating in our detection challenge
0137
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 19/09/2026
The [DRAFT] program for AVSP 2026 is out! Check out project.inria.fr/avsp2026/wor... A terrific day to be had: Great presentations throughout An excellent keynote project.inria.fr/avsp2026/key... A 50th anniversary event re. the McGurk & MacDonald effect (there may be cake🎂) A lipreading contest😦
project.inria.fr
Workshop program – AVSP2026
032
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 18/09/2026
Multimodal Communication in Autism Is Reorganized, Not Reduced: Evidence From Speech and Gesture Across Four Languages journals.sagepub.com/doi/10.1177/... Autistic & non-autistic children tell stories, speech & gesture analyzed wrt amount, diversity & complexity.
journals.sagepub.com
Multimodal Communication in Autism Is Reorganized, Not Reduced: Evidence From Speech and Gesture Across Four Languages - Armita Ghobadi, Pauline Wolfer, Stephanie Durrleman, 2026
Speech and gesture form an integrated communicative system, yet how this system operates in autism, and whether it varies across languages, remains unclear. We ...
010
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 18/09/2026
Exploring the acceptability of a lip-reading app in acute and critical care: Perspectives from patients with tracheostomies, relatives and healthcare professional www.sciencedirect.com/science/arti... Study aimed to explore the acceptability of SRAVI among tracheostomies patients relatives & staff
sciencedirect.com
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 18/09/2026
Multimodal Language Processing in Children With Developmental Language Disorder: The Influence of Verbal and Visuospatial Memory pubs.asha.org/doi/10.1044/... How do verbal & visuospatial memory capacities contribute to the processing & comprehension of pragmatic meanings in children with/out DLD?
pubs.asha.org
Multimodal Language Processing in Children With Developmental Language Disorder: The Influence of Verbal and Visuospatial Memory
Purpose: This study investigates how verbal and visuospatial memory capacities contribute to the processing and comprehension of pragmatic meanin...
000
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 17/09/2026
📚 Citation Classic "Robust text-independent speaker identification using Gaussian mixture speaker models" Reynolds & Rose (1995) Citations: 7600+ GMM-based... 🔗 scholar.google.com/scholar?q=Reynol… #SpeechScience
031
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 16/09/2026
Larger language models better align with neural representations of natural language elifesciences.org/articles/101... Fitted electrode-wise (ECoG) encoding models from contextual embeddings of each hidden layer of LLMs to predict word-level neural signals. The early layers of larger LLMs do better..
elifesciences.org
Larger language models better align with neural representations of natural language
Embeddings from larger language models and relatively earlier layers better predict neural activity during natural language.
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 16/09/2026
Simultaneous speech and gesture decoding for multimodal communication in paralysis www.nature.com/articles/s41... Shows neural signals recorded with a single high-density electrocorticography implant can support simultaneous decoding of speech & gestures Commentary www.nature.com/articles/d41...
nature.com
Simultaneous speech and gesture decoding for multimodal communication in paralysis - Nature Neuroscience
Brosler et al. develop a brain–computer interface that simultaneously decodes speech and gestures from a single cortical implant to animate a virtual avatar, providing a step toward more flexible comm...
061
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 13/09/2026
Q/ Do visual cues disambiguate fine-grained articulatory features during early perceptual stages or integrate with speech at phoneme-level stages? A/ AV speech ⬆️ classifier confidence & decoding accuracy for phoneme-level representations without corresponding effects on phonetic features
010
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 13/09/2026
Visual speech enhances phoneme separability in human superior temporal gyrus www.biorxiv.org/content/10.64898/20…
121
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 12/09/2026
Impact of visual mouth image presentation on speech perception in noise in normal hearing subjects and cochlear implant users journals.plos.org/plosone/arti... Computer‑animated avatar provided no speech perception AV benefit; real‑speaker videos enhanced AV speech perception but no impact on effort
journals.plos.org
Impact of visual mouth image presentation on speech perception in noise in normal hearing subjects and cochlear implant users
This study investigated the impact of visual speech cues on speech perception in noise and subjective listening effort in cochlear implant (CI) users compared with normal-hearing (NH) individuals. Fur...
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 11/09/2026
The Cortical Contribution to the Speech-Frequency–Following Response Is Not Modulated by Visual Information onlinelibrary.wiley.com/doi/10.1111/... "Our results suggest that visual modulation of the speech-FFR in the auditory cortex is, if existent, too small to be measurable"
onlinelibrary.wiley.com
001
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 04/09/2026
Next time when you're thinking of zooming in on that meeting when driving ... "Fast cars, slow words? Tracking processing of concurrent speaking and driving in real time" link.springer.com/article/10.3...
link.springer.com
Fast cars, slow words? Tracking processing of concurrent speaking and driving in real time - Psychonomic Bulletin & Review
The study of multitasking provides a unique window into the interaction of cognitive processes and systems and as such, into the cognitive architecture of the human mind. Previous studies reported tha...
030
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 03/09/2026
The shrinking landscape of linguistic diversity in the age of large language models www.nature.com/articles/s41... Egads, the trumpet tongues of scholars now sound a single tune; gone, the curate's egg of runny then congealed prose, all is pabulum.
nature.com
The shrinking landscape of linguistic diversity in the age of large language models - Nature Human Behaviour
Sourati et al. find that variance in writing complexity dropped on social media, news and scientific writing after ChatGPT’s release, a shift from individuality towards uniformity. LLM polishing strip...
032
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 01/09/2026
Rhythm Across Species: Situating Human Speech in a Comparative Framework nyaspubs.onlinelibrary.wiley.com/doi/10.1111/... Examined speech rhythm using units used for nonhuman animals: interonset intervals (IOI) found these durations are relatively similar across a sample of 48 languages - we fit in🥁
nyaspubs.onlinelibrary.wiley.com
NYAS Publications
We analyzed speech rhythms from 48 human languages using methods drawn from biology and linguistics. Interonset intervals across languages are surprisingly similar, with a median value of 2 seconds.,...
031
Reposted by Auditory-Visual Speech Association (AVISA)
Ev Fedorenko @evfedorenko.bsky.social · 01/09/2026
Go, @andreadevarda.bsky.social! A beautiful and comprehensive study! 🔑 findng: behav. measures are ~fully reducible to simple predictors of processing effort (surprisal, word length+frequency), but for 🧠 measures, LLM embeddings carry additional predictive power, likely capturing aspects of meaning.
094
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 30/08/2026
The role of high-amplitude bursts of high-gamma activity in naturalistic speech and music listening direct.mit.edu/imag/article... Neuronal avalanches (top 1% of high-gamma activations) may provide an insight into stimulus-relevant, large-scale neural dynamics - its always the bursty noise makers
direct.mit.edu
The role of high-amplitude bursts of high-gamma activity in naturalistic speech and music listening
Abstract. A central challenge in systems neuroscience is to understand how distributed brain networks organize activity during naturalistic cognition. Here, we investigate whether transient, high-ampl...
041
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 29/08/2026
Evaluating perceptual evidence of cross-modal plasticity in older adults with hearing loss using an AV word mismatch task www.sciencedirect.com/science/arti... Audio/visual mismatch detection task noise/no noise; greater hearing loss not associated with visual benefits or altered gaze to face
sciencedirect.com
010
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 27/08/2026
📚 Citation Classic "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations" Baevski et al (2020) Citations: 9100+ wav2vec 2.0 showed self-supervised pretraining slashes the need for labeled speech. 🔗 arxiv.org/abs/2006.11477 #SpeechScience
121
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 27/08/2026
Gee, Wav2Vec [1] Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised pre-training for speech recognition arxiv.org/abs/1904.05862 only has 2371 cites
arxiv.org
wav2vec: Unsupervised Pre-training for Speech Recognition
We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are ...
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 27/08/2026
Sensory context improves language prediction in humans & LLMs www.pnas.org/doi/abs/10.1... Compared humans predicting language in reading, listening, watching AV videos with text-based LLM predictions. People do better, AV ~ AO > reading (some of it is prosody); LMMs do better when given this info👁️👂
pnas.org
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
010
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 23/08/2026
pure.mpg.de/rest/items/i... “This paper describes a number of objective experiments on recognition, concerning particularly the relation between the messages received by the two ears. Rather than use steady tones or clicks (frequency or time-point signals) continuous speech is used...
pure.mpg.de
000
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 23/08/2026
Colin Cherry (1953) described how we can focus on one conversation in a noisy room full of people talking. This launched decades of research into selective attention and speech perception. 🎧 Cherry (1953) - foundational auditory perception study #SpeechScience #Perception
142
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 23/08/2026
Assessing Visual Speech Communication Abilities: Cross-Situational Consistency in Lipreadability brill.com/view/journal... Visual speech abilities: Receivers consistently good or poor at lipreading regardless of the talker;OR, talkers consistently lipread accurately or not, regardless of listener 👁️😮
brill.com
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 21/08/2026
Viewpoint on Multisensory Time Perception brill.com/view/journal... Shi & van Wassenhove draw on decades of research & argue that the field has undergone an important shift from the centralized internal clock toward timing as an active inference process.
brill.com
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 21/08/2026
I think the link should be dl.acm.org/doi/pdf/10.1... Your link is to Graves, A. (2012). Connectionist temporal classification. In Supervised sequence labelling with recurrent neural networks (pp. 61-93). Berlin, Heidelberg: Springer Berlin Heidelberg.
dl.acm.org
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 20/08/2026
Voice perception reflects graded sensitivity to acoustic cues along the auditory hierarchy www.sciencedirect.com/science/arti... How like a human voice? Acoustic patterns over time; Perceptual & neural features align most in the auditory association area; voiceness: a hierarchical encoding framework
sciencedirect.com
Voice perception reflects graded sensitivity to acoustic cues along the auditory hierarchy
Humans readily recognize voices in noisy, challenging environments, yet how the brain does this remains poorly understood. A key question is whether t…
021
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 19/08/2026
Music Is Quasi-Rhythmic Too: Comparing Bandwidth in Speech and Music Acoustic Modulation Spectra. nyaspubs.onlinelibrary.wiley.com/doi/10.1111/... Confirms findings that speech & music modulation spectra differ in their centre frequency -> but find that the modulation bandwidth is comparable 🎵 ↔️ 💬
nyaspubs.onlinelibrary.wiley.com
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 16/08/2026
Ok - you just know I'll favour this - even tho you've posted it before ...🫢
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 15/08/2026
Distinct neural processes link speech planning and execution www.nature.com/articles/s41... Busy day on the intracranial recordings research front: planning => prefrontal units => dynamically integrated => execution via speech motor regions representing speech sound units & transitional properties
nature.com
Distinct neural processes link speech planning and execution - Nature Human Behaviour
Direct recordings from the human brain reveal a hierarchy in planning to speak, with neural activity organizing whole syllables before the individual sounds that compose them are sequenced.
021
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 14/08/2026
Invariant neural representation of parts of speech in the human brain www.cell.com/current-biol... 1,801 electrodes,20 patient, 2 word adjective noun phrases, representation of parts of speech invariant across isual and auditory presentation modalities! Robust to word order, length, frequency ...
cell.com
Invariant neural representation of parts of speech in the human brain
Misra et al. use invasive neurophysiological recordings from the human brain to describe a highly circumscribed region within the left lateral orbitofrontal cortex that shows strong selectivity for pa...
040
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 14/08/2026
Signal-specific performance of in-ear EEG: strengths and limitations www.frontiersin.org/journals/neu... Results indicate signal detectability depends strongly on signal class, effect size, & recording geometry ->so, a configuration-aware guidance for study design and future in-ear EEG development 👍
frontiersin.org
Frontiers | Signal-specific performance of in-ear EEG: strengths and limitations
Fully in-ear Electroencephalography (EEG) configurations prioritize wearability and rapid setup but may constrain spatial sampling and signal-to-noise charac...
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 13/08/2026
Audiovisual speech perception in Cantonese-speaking children: Developmental trajectories and the role of noise academic.oup.com/chidev/advan... Results suggest that children gradually learn to use the most reliable source of information depending on listening conditions
academic.oup.com
Audiovisual speech perception in Cantonese-speaking children: Developmental trajectories and the role of noise
Abstract. The development of audiovisual speech perception in tone-language-speaking children remains debated, and this study addressed this issue by exami
000
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 13/08/2026
Real-Time Conversation Comprehension in Younger and Older Adults: The Effects of Noise, Visual Speech, Speaker Switch, and Context journals.sagepub.com/doi/10.1177/... A fun one to put together, we used a topic monitoring task to assess real-time semantic processing, we think this task is useful! 👍
journals.sagepub.com
Real-Time Conversation Comprehension in Younger and Older Adults: The Effects of Noise, Visual Speech, Speaker Switch, and Context - Chris Davis, April Ching, Afrah Haroon, Jeesun Kim, 2026
This study investigated older and younger adults’ conversational speech understanding using an established topic monitoring task to assess real-time semantic pr...
030
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 11/08/2026
Pedagogy in the speech-gesture couplings of caregivers: Evidence from a corpus-based analysis pubmed.ncbi.nlm.nih.gov/42574797
011
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 09/08/2026
Atypical low-frequency & high-frequency neural entrainment to rhythmic audiovisual speech in adults with dyslexia direct.mit.edu/imag/article... EEG study: 24 dyslexic & 24 control adults presented with a “talking head” repeating the syllable “ba” at 2-Hz entrainment & band power measured >see title
direct.mit.edu
Atypical low-frequency and high-frequency neural entrainment to rhythmic audiovisual speech in adults with dyslexia
Abstract. Developmental dyslexia has been linked to atypical neural processing of the temporal dynamics of speech, but there has been disagreement concerning whether faster or slower dynamics are impa...
020
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 09/08/2026
In 1961, physicist John Kelly programmed an IBM 704 to sing 'Daisy Bell' - the first song ever sung by a computer. This inspired HAL 9000's song in 2001: A Space Odyssey! 🎵 Historic: youtube.com/watch?v=41U78QP8nBk #SpeechScience #Technology
011
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 08/08/2026
Eye movements reveal a dissociation between prediction and structural processing difficulty in language comprehension www.pnas.org/doi/abs/10.1... Surprisal (409 estimates from language models with multiple architectures & training settings) explains early not late stages of reading disambiguation
pnas.org
Eye movements reveal a dissociation between prediction and structural processing difficulty in language comprehension | PNAS
In the process of extracting a meaning from a text, our eyes linger much more on some words than others, and we often reread earlier portions of th...
120
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 07/08/2026
When rhythm guides speech production: Insights from Parkinson's disease and spinocerebellar ataxia www.sciencedirect.com/science/arti... Do brief rhythmic primes modulate speech production & facilitate it in people with Parkinson's disease & spinocerebellar ataxia?
sciencedirect.com
031
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 06/08/2026
📚 Citation Classic "Robust Speech Recognition via Large-Scale Weak Supervision" Radford et al (2022) Citations: 9600+ Whisper's weak supervision at scale made robust multilingual ASR go mainstream. 🔗 arxiv.org/abs/2212.04356 #SpeechScience
111
Reposted by Auditory-Visual Speech Association (AVISA)
Virginie van Wassenhove @virginievanw.bsky.social · 06/08/2026
A very interesting coverage of the recent work from the lab www.nature.com/articles/s41... conducted by @matthewrlogie.bsky.social & @grassocamille.bsky.social @brainthemind.bsky.social @unicog.bsky.social
nature.com
Nested contextual change and the temporal compression of episodic memory - Scientific Reports
Scientific Reports - Nested contextual change and the temporal compression of episodic memory
02410
Reposted by Auditory-Visual Speech Association (AVISA)
speechpapers.bsky.social @speechpapers.bsky.social · 05/08/2026
Is that clear? Robust electrophysiological measures of the effects of prior knowledge on degraded speech perception. www.biorxiv.org/content/10.64898/20…
032
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 04/08/2026
for comment see ¹ ¹ See ²
010