Simplified diarization pipeline using some pretrained models - audio file to diarized segments in a few lines of code

Last update: Dec 30, 2022

Overview

simple_diarizer

Simplified diarization pipeline using some pretrained models.

Made to be a simple as possible to go from an input audio file to diarized segments.

import soundfile as sf
import matplotlib.pyplot as plt

from simple_diarizer.diarizer import Diarizer
from simple_diarizer.utils import combined_waveplot

diar = Diarizer(
                  embed_model='xvec', # 'xvec' and 'ecapa' supported
                  cluster_method='sc' # 'ahc' and 'sc' supported
               )

segments = diar.diarize(WAV_FILE, num_speakers=NUM_SPEAKERS)

signal, fs = sf.read(WAV_FILE)
combined_waveplot(signal, fs, segments)
plt.show()

Source Video

"Some Quick Advice from Barack Obama!"

Pre-trained Models

The following pretrained models are used:

Voice Activity Detection (VAD)
- Silero VAD
Deep speaker embedding extraction
- SpeechBrain
  - X-Vector
  - ECAPA-TDNN
(Optional/Experimental) Speech-to-text
- ESPnet Model Zoo
  - English ASR model

Demo

It can be checked out in the above link, where it will try and diarize any input YouTube URL. It will also use YouTube's autogenerated transcriptions to produce a speaker labelled transcription.

Hopefully this can be of use as a free basic tool to produce a diarized transcript of a video/audio of interest.

Other References

Spectral clustering methods lifted from https://github.com/wq2012/SpectralCluster

Planned Features

Comments

WIP - Make an installable package

Description:

Include requirements.txt.
Add setup*. files to build a package.
Create a folder simple_diarizer to store source code.
Create Github Workflow to publish the package.

How to test:

Run command pip install .
Outside project folder type python and from simple_diarizer import diarizer

Notes:

Cannot use python 3.10.x yet

Source code to test:

from simple_diarizer.utils import (convert_wavfile, download_youtube_wav)

from simple_diarizer.diarizer import Diarizer
import tempfile

YOUTUBE_ID = "HyKmkLEtQbs"

with tempfile.TemporaryDirectory() as outdir:
    yt_file = download_youtube_wav(YOUTUBE_ID, outdir)

    wav_file = convert_wavfile(yt_file, f"{outdir}/{YOUTUBE_ID}_converted.wav")

    print(f"wav file: {wav_file}")

    diar = Diarizer(
        embed_model='ecapa', # supported types: ['xvec', 'ecapa']
        cluster_method='sc', # supported types: ['ahc', 'sc']
        window=1.5, # size of window to extract embeddings (in seconds)
        period=0.75 # hop of window (in seconds)
    )

    NUM_SPEAKERS = 2

    segments = diar.diarize(wav_file, 
                            num_speakers=NUM_SPEAKERS,
                            outfile=f"{outdir}/{YOUTUBE_ID}.rttm")

    print(segments)

opened by johnidm 16

"[Errno 30] Read-only file system: 'pretrained_models'"

I am using macOS and I am getting error "[Errno 30] Read-only file system: 'pretrained_models'" From what I can tell, the pretrained models are being fetched if you do not have them.

However, the save location is the root directory which is read-only. This is where I believe is the target directory "./pretrained_model_checkpoints"

Is there another location that can be used that can be used?

PythonKit/Python.swift:706: Fatal error: 'try!' expression unexpectedly raised an error: Python exception: [Errno 30] Read-only file system: 'pretrained_models' Traceback: File "/Users/wedwards/Documents/Development/A_PythonKit_Test/A_PythonKit_Test/Simple Diarizer.py", line 42, in diar = Diarizer( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/simple_diarizer/diarizer.py", line 48, in init self.embed_model = EncoderClassifier.from_hparams( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/speechbrain/pretrained/interfaces.py", line 342, in from_hparams hparams_local_path = fetch( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/speechbrain/pretrained/fetching.py", line 86, in fetch savedir.mkdir(parents=True, exist_ok=True) File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/pathlib.py", line 1179, in mkdir self.parent.mkdir(parents=True, exist_ok=True) File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/pathlib.py", line 1175, in mkdir self._accessor.mkdir(self, mode)

2022-11-11 13:14:00.531470-0500 A_PythonKit_Test[69382:7584330] PythonKit/Python.swift:706: Fatal error: 'try!' expression unexpectedly raised an error: Python exception: [Errno 30] Read-only file system: 'pretrained_models' Traceback: File "/Users/wedwards/Documents/Development/A_PythonKit_Test/A_PythonKit_Test/Simple Diarizer.py", line 42, in diar = Diarizer( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/simple_diarizer/diarizer.py", line 48, in init self.embed_model = EncoderClassifier.from_hparams( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/speechbrain/pretrained/interfaces.py", line 342, in from_hparams hparams_local_path = fetch( File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/site-packages/speechbrain/pretrained/fetching.py", line 86, in fetch savedir.mkdir(parents=True, exist_ok=True) File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/pathlib.py", line 1179, in mkdir self.parent.mkdir(parents=True, exist_ok=True) File "/Library/Frameworks/Python.framework/Versions/3.10/lib/python3.10/pathlib.py", line 1175, in mkdir self._accessor.mkdir(self, mode)

opened by MrEdwards007 5
Latest Python and packages

The current release prevents use of Python 3.10 and requires specific versions of Beautiful Soup and PyTube.

I've forked the repo to overcome these version limitations and it's working for me. I haven't made a pull request, however, as your repo doesn't have tests and I don't know whether there is a use case which would be broken by my changes.

Can you please remove these version limitations if they're not needed?

Thanks for the repo - it's effective and much easier to use than SpeechBrain.

opened by andrewmackie 3
takes 1 positional argument but 2 were given

running a demo on google co-lab i am getting the following error, any idea how to resolve this,

File "/root/anaconda3/envs/simple/lib/python3.8/site-packages/speechbrain/pretrained/fetching.py", line 116, in fetch fetched_file = huggingface_hub.cached_download(url, use_auth_token) TypeError: cached_download() takes 1 positional argument but 2 were given

opened by SanaullahOfficial 2

AttributeError when running Diarizer in simple_diarizer.diarizer

Hi there!

When running the following code in Python 3.7 on a fresh conda environment in Ubuntu 22.04

from simple_diarizer.diarizer import Diarizer

diar = Diarizer(
                    embed_model='xvec', # 'xvec' and 'ecapa' suported
                    cluster_method='sc' # 'ahc' and 'sc' supported
                )

I get the following error:

<ipython-input-3-286690ce0195> in <module>
      1 diar = Diarizer(
      2                     embed_model='xvec', # 'xvec' and 'ecapa' suported
----> 3                     cluster_method='sc' # 'ahc' and 'sc' supported
      4                 )

~/anaconda3/envs/test/lib/python3.7/site-packages/simple_diarizer/diarizer.py in __init__(self, embed_model, cluster_method, window, period)
     44             self.embed_model = EncoderClassifier.from_hparams(source="speechbrain/spkrec-xvect-voxceleb",
     45                                                               savedir="pretrained_models/spkrec-xvect-voxceleb",
---> 46                                                               run_opts=self.run_opts)
     47         if embed_model == 'ecapa':
     48             self.embed_model = EncoderClassifier.from_hparams(source="speechbrain/spkrec-ecapa-voxceleb",

~/anaconda3/envs/test/lib/python3.7/site-packages/speechbrain/pretrained/interfaces.py in from_hparams(cls, source, hparams_file, pymodule_file, overrides, savedir, use_auth_token, **kwargs)
    349         # Load the modules:
    350         with open(hparams_local_path) as fin:
--> 351             hparams = load_hyperpyyaml(fin, overrides)
    352 
    353         # Pretraining:

~/anaconda3/envs/test/lib/python3.7/site-packages/hyperpyyaml/core.py in load_hyperpyyaml(yaml_stream, overrides, overrides_must_match)
    187 
    188     # Remove items that start with "__"
--> 189     removal_keys = [k for k in hparams.keys() if k.startswith("__")]
    190     for key in removal_keys:
    191         del hparams[key]

AttributeError: 'str' object has no attribute 'keys'

opened by masonhargrave 2

Make project installable
Hi @cvqluu, this project is amazing, thanks for sharing.

I have some experience in packaging projects in Python.

What do you think I make these items on your to-do list?

Add to PyPi (make pip installable)

requirements.txt

If you authorize me, I will start doing this now and submit pull requests for your review and approval.
opened by johnidm 1
Added ipython depedency
Tested on local machine using:

pip install --user git+https://github.com/cvqluu/[email protected]

Fix for https://github.com/cvqluu/simple_diarizer/issues/12
opened by cvqluu 0
Bump ipython from 7.30.1 to 7.31.1
Bumps ipython from 7.30.1 to 7.31.1.

Commits

e321e76 release 7.31.1

67ca2b3 Merge pull request from GHSA-pq7m-3gw7-gq5x

2794330 back to dev

be343e7 release 7.31.0

0fcf2c4 Merge pull request #13428 from meeseeksmachine/auto-backport-of-pr-13427-on-7.x

b8db9b1 Backport PR #13427: wn 731

7f253dc Merge pull request #13412 from bnavigator/backport-inspect

4f26796 fix xxlimited_35 import name

77ca4a6 don't run nose-based iptest on py310, only pytest

533e509 back to decorator skip

Additional commits viewable in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.

Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

@dependabot rebase will rebase this PR

@dependabot recreate will recreate this PR, overwriting any edits that have been made to it

@dependabot merge will merge this PR after your CI passes on it

@dependabot squash and merge will squash and merge this PR after your CI passes on it

@dependabot cancel merge will cancel a previously requested merge and block automerging

@dependabot reopen will reopen this PR if it is closed

@dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually

@dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)

@dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)

@dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

@dependabot use these labels will set the current labels as the default for future PRs for this repo and language

@dependabot use these reviewers will set the current reviewers as the default for future PRs for this repo and language

@dependabot use these assignees will set the current assignees as the default for future PRs for this repo and language

@dependabot use this milestone will set the current milestone as the default for future PRs for this repo and language

You can disable automated security fix PRs for this repo from the Security Alerts page.

dependencies
opened by dependabot[bot] 0

Undeclared IPython dependency

The current package (0.0.12 on PyPI) cannot run without IPython, but this is missing from requirements.txt

Steps to reproduce (outside of a Jupyter notebook):

pip install simple-diarizer

# index.py
from simple_diarizer.diarizer import Diarizer

Output:

File "[redacted]\index.py", line 1, in <module>
    from simple_diarizer.diarizer import Diarizer
File "[redacted]\lib\site-packages\simple_diarizer\diarizer.py", line 13, in <module>
    from .utils import check_wav_16khz_mono, convert_wavfile
File "[redacted]\lib\site-packages\simple_diarizer\utils.py", line 8, in <module>
    from IPython.display import Audio, display
ModuleNotFoundError: No module named 'IPython'

opened by DavidRalph 1

waveplot_perspeaker causes argument out of range error

While running through your code example, testing the workflow on a different audio file produced the following output:

C:\Users\xxx\Miniconda3\envs\simple_diarizer_env\lib\site-packages\IPython\lib\display.py:187: RuntimeWarning: invalid value encountered in divide
  scaled = data / normalization_factor * 32767
---------------------------------------------------------------------------
error                                     Traceback (most recent call last)
Cell In [18], line 1
----> 1 waveplot_perspeaker(signal, fs, segments)

File ~\Miniconda3\envs\simple_diarizer_env\lib\site-packages\simple_diarizer\utils.py:166, in waveplot_perspeaker(signal, fs, segments)
    164 if "words" in seg:
    165     pprint(seg["words"])
--> 166 display(Audio(speech, rate=fs))
    167 print("=" * 40 + "\n")

File ~\Miniconda3\envs\simple_diarizer_env\lib\site-packages\IPython\lib\display.py:130, in Audio.__init__(self, data, filename, url, embed, rate, autoplay, normalize, element_id)
    128 if rate is None:
    129     raise ValueError("rate must be specified when data is a numpy array or list of audio samples.")
--> 130 self.data = Audio._make_wav(data, rate, normalize)

File ~\Miniconda3\envs\simple_diarizer_env\lib\site-packages\IPython\lib\display.py:162, in Audio._make_wav(data, rate, normalize)
    160 waveobj.setsampwidth(2)
    161 waveobj.setcomptype('NONE','NONE')
--> 162 waveobj.writeframes(scaled)
    163 val = fp.getvalue()
    164 waveobj.close()

File ~\Miniconda3\envs\simple_diarizer_env\lib\wave.py:437, in Wave_write.writeframes(self, data)
    436 def writeframes(self, data):
--> 437     self.writeframesraw(data)
    438     if self._datalength != self._datawritten:
    439         self._patchheader()

File ~\Miniconda3\envs\simple_diarizer_env\lib\wave.py:426, in Wave_write.writeframesraw(self, data)
    424 if not isinstance(data, (bytes, bytearray)):
    425     data = memoryview(data).cast('B')
--> 426 self._ensure_header_written(len(data))
    427 nframes = len(data) // (self._sampwidth * self._nchannels)
    428 if self._convert:

File ~\Miniconda3\envs\simple_diarizer_env\lib\wave.py:467, in Wave_write._ensure_header_written(self, datasize)
    465 if not self._framerate:
    466     raise Error('sampling rate not specified')
--> 467 self._write_header(datasize)

File ~\Miniconda3\envs\simple_diarizer_env\lib\wave.py:479, in Wave_write._write_header(self, initlength)
    477 except (AttributeError, OSError):
    478     self._form_length_pos = None
--> 479 self._file.write(struct.pack('<L4s4sLHHLLHH4s',
    480     36 + self._datalength, b'WAVE', b'fmt ', 16,
    481     WAVE_FORMAT_PCM, self._nchannels, self._framerate,
    482     self._nchannels * self._framerate * self._sampwidth,
    483     self._nchannels * self._sampwidth,
    484     self._sampwidth * 8, b'data'))
    485 if self._form_length_pos is not None:
    486     self._data_length_pos = self._file.tell()

error: argument out of range

Any ideas what the issue could be? It works fine on other audio files, and everything up to this point seems to run without error.

opened by dcruiz01 1

Releases(v0.0.13)

v0.0.13(Dec 12, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.12(Dec 8, 2022)

Setting extra_info to True will now return an additional dict, containing cluster labels
Source code(tar.gz)
Source code(zip)
v0.0.11(Nov 9, 2022)

Removed youtube related dependencies, keeping the repository slim. There are no longer youtube helper functions, but the core functionality should now work for python >=3.7
Source code(tar.gz)
Source code(zip)
v0.0.10(Aug 30, 2022)

Allowed for a newer version of speechbrain, which should have fixed the issues with pulling from huggingface_hub
Source code(tar.gz)
Source code(zip)
v0.0.9(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.8(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.7(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.6(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.5(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.4(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.3(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.2(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)
v0.0.1(Jan 10, 2022)

Source code(tar.gz)
Source code(zip)

Owner

Chau

PhD student at the University of Edinburgh, CSTR

GitHub Repository

This repository contains the code for EMNLP-2021 paper "Word-Level Coreference Resolution"

Word-Level Coreference Resolution This is a repository with the code to reproduce the experiments described in the paper of the same name, which was a

79 Dec 27, 2022

MRC approach for Aspect-based Sentiment Analysis (ABSA)

B-MRC MRC approach for Aspect-based Sentiment Analysis (ABSA) Paper: Bidirectional Machine Reading Comprehension for Aspect Sentiment Triplet Extracti

1 Apr 05, 2022

Final Project for the Intel AI Readiness Boot Camp NLP (Jan)

NLP Boot Camp (Jan) Synopsis Full Name: Prameya Mohanty Name of your School: Delhi Public School, Rourkela Class: VIII Title of the Project: iTransect

1 Feb 01, 2022

Code for the Findings of NAACL 2022(Long Paper): AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks

AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks arXiv link: upcoming To be published in Findings of NA

16 Nov 12, 2022

ChainKnowledgeGraph, 产业链知识图谱包括A股上市公司、行业和产品共3类实体

ChainKnowledgeGraph, 产业链知识图谱包括A股上市公司、行业和产品共3类实体，包括上市公司所属行业关系、行业上级关系、产品上游原材料关系、产品下游产品关系、公司主营产品、产品小类共6大类。上市公司4,654家，行业511个，产品95,559条、上游材料56,824条，上级行业480条，下游产品390条，产品小类52,937条，所属行业3,946条。

415 Jan 06, 2023

Official source for spanish Language Models and resources made @ BSC-TEMU within the "Plan de las Tecnologías del Lenguaje" (Plan-TL).

Spanish Language Models 💃🏻 A repository part of the MarIA project. Corpora 📃 Corpora Number of documents Number of tokens Size (GB) BNE 201,080,084

203 Dec 20, 2022

In this repository, I have developed an end to end Automatic speech recognition project. I have developed the neural network model for automatic speech recognition with PyTorch and used MLflow to manage the ML lifecycle, including experimentation, reproducibility, deployment, and a central model registry.

End to End Automatic Speech Recognition In this repository, I have developed an end to end Automatic speech recognition project. I have developed the

22 Nov 13, 2022

Auto-researching tool generating word documents.

About ResearchTE automates researching by generating document with answers to given questions. Supports getting results from: Google DuckDuckGo (with

1 Feb 14, 2022

Ceaser-Cipher - The Caesar Cipher technique is one of the earliest and simplest method of encryption technique

Ceaser-Cipher The Caesar Cipher technique is one of the earliest and simplest me

2 May 12, 2022

Weakly-supervised Text Classification Based on Keyword Graph

Weakly-supervised Text Classification Based on Keyword Graph How to run? Download data Our dataset follows previous works. For long texts, we follow C

20 Dec 29, 2022

Unofficial Python library for using the Polish Wordnet (plWordNet / Słowosieć)

Polish Wordnet Python library Simple, easy-to-use and reasonably fast library for using the Słowosieć (also known as PlWordNet) - a lexico-semantic da

12 Dec 23, 2022

Gold standard corpus annotated with verb-preverb connections for Hungarian.

Hungarian Preverb Corpus A gold standard corpus manually annotated with verb-preverb connections for Hungarian. corpus The corpus consist of the follo

3 Jan 27, 2022

End-to-End Speech Processing Toolkit

ESPnet: end-to-end speech processing toolkit system/pytorch ver. 1.0.1 1.1.0 1.2.0 1.3.1 1.4.0 1.5.1 1.6.0 1.7.1 1.8.1 ubuntu18/python3.8/pip ubuntu18

5.9k Jan 03, 2023

Utilizing RBERT model for KLUE Relation Extraction task

RBERT for Relation Extraction task for KLUE Project Description Relation Extraction task is one of the task of Korean Language Understanding Evaluatio

14 Nov 15, 2022

This simple Python program calculates a love score based on your and your crush's full names in English

This simple Python program calculates a love score based on your and your crush's full names in English. There is no logic or reason in the calculation behind the love score. The calculation could ha

1 Jan 24, 2022

chaii - hindi & tamil question answering

chaii - hindi & tamil question answering This is the solution for rank 5th in Kaggle competition: chaii - Hindi and Tamil Question Answering. The comp

33 Dec 18, 2022

Uses Google's gTTS module to easily create robo text readin' on command.

Tool to convert text to speech, creating files for later use. TTRS uses Google's gTTS module to easily create robo text readin' on command.

0 Jun 20, 2021

A python script that will use hydra to get user and password to login to ssh, ftp, and telnet

Hydra-Auto-Hack A python script that will use hydra to get user and password to login to ssh, ftp, and telnet Project Description This python script w

2 Jan 16, 2022

A Fast Sequence Transducer Implementation with PyTorch Bindings

transducer A Fast Sequence Transducer Implementation with PyTorch Bindings. The corresponding publication is Sequence Transduction with Recurrent Neur

184 Dec 18, 2022

A fast Text-to-Speech (TTS) model. Work well for English, Mandarin/Chinese, Japanese, Korean, Russian and Tibetan (so far). 快速语音合成模型，适用于英语、普通话/中文、日语、韩语、俄语和藏语（当前已测试）。

简体中文 | English 并行语音合成 [TOC] 新进展 2021/04/20 合并 wavegan 分支到 main 主分支，删除 wavegan 分支！ 2021/04/13 创建 encoder 分支用于开发语音风格迁移模块！ 2021/04/13 softdtw 分支支持使用 Sof

161 Dec 19, 2022