-
phonespam Public
Phone segmentation and recognition through S3M-based Phonological Activation Mapping (SPAM).
-
epitran Public
Forked from dmort27/epitranA tool for transcribing orthographic text as IPA (International Phonetic Alphabet)
-
prism Public
Forked from changelinglab/prismA toolkit and benchmark for evaluating phonetic capabilities of speech models.
Python UpdatedJun 10, 2026 -
-
specplotter Public
Python library for nice-looking spectrograms.
-
-
vocos Public
Forked from gemelo-ai/vocosVocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
-
unbox-w2v-convnet Public
Official implementation of the paper, "Opening the Black Box of wav2vec Feature Encoder."
-
changelinglab.github.io Public
Forked from changelinglab/changelinglab.github.ioWebsite for the CMU Language Change and Empirical Linguistics Lab
HTML UpdatedOct 7, 2025 -
-
espnet Public
Forked from espnet/espnetEnd-to-End Speech Processing Toolkit
-
starter-repo Public
Forked from neubig/starter-repoAn example starter repo for Python projects
Python MIT License UpdatedJun 16, 2025 -
slamkit Public
Forked from slp-rl/slamkitSlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"
Python MIT License UpdatedMay 18, 2025 -
acoustic-units-for-ood Public
Official implementation for the paper "Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment (NAACL 2025)"
-
dysarthria-gop Public
Official implementation of the paper "Speech Intelligibility Assessment of Dysarthric Speech by using Goodness of Pronunciation with Uncertainty Quantification" (Interspeech 2023)
-
neural-transducer Public
Forked from shijie-wu/neural-transducerThis repo contains a set of neural transducer, e.g. sequence-to-sequence model, focusing on character-level tasks.
-
-
-
-
shinjiwlab.github.io Public
Forked from wavlab-speech/shinjiwlab.github.ioJavaScript MIT License UpdatedOct 22, 2024 -
phonetic_semantic_probing Public
Official implementation of the paper "Self-Supervised Speech Representations are More Phonetic than Semantic (Interspeech 2024)"
-
-
ldc_downloader Public
Forked from dowobeha/ldc_downloaderScript to download corpora from the Linguistic Data Consortium (LDC)
-
dynamic-superb Public
Forked from dynamic-superb/dynamic-superbThe official repository of Dynamic-SUPERB.
Python UpdatedAug 5, 2024 -
information_probing Public
Official implementation of the paper "Understanding Probe Behaviors through Variational Bounds of Mutual Information (ICASSP 2024)"
-
KILT Public
Forked from facebookresearch/KILTLibrary for Knowledge Intensive Language Tasks
Python MIT License UpdatedMay 2, 2024 -
xlm_to_xlsr Public
Official implementation of the paper "Distilling a Pretrained Language Model to a Multilingual ASR Model" (Interspeech 2022)
-
ZMM-TTS Public
Forked from nii-yamagishilab/ZMM-TTSZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
C BSD 3-Clause "New" or "Revised" License UpdatedMar 6, 2024 -
panphon Public
Forked from dmort27/panphonPython package and data files for manipulating phonological segments (phones, phonemes) in terms of universal phonological features.
Python MIT License UpdatedFeb 29, 2024 -
dysarthria-mtl Public
Official implementation of the paper "Automatic Severity Assessment of Dysarthric speech by using Self-supervised Model with Multi-task Learning (ICASSP 2023)"






