我有电话交谈的录音,
我使用了Simulareyzer,它根据扬声器对音频进行聚类。输出是labelling
,这基本上是一个字典,它记录了人说话的时间(说话人标签、开始时间、结束时间)
我需要分段音频输出扬声器明智的基础上的时间在标签。我已经做了一个星期了
from resemblyzer import preprocess_wav, VoiceEncoder
from pathlib import Path
import pickle
import scipy.io.wavfile
from spectralcluster import SpectralClusterer
audio_file_path = 'C:/Users/...'
wav_fpath = Path(audio_file_path)
wav = preprocess_wav(wav_fpath)
encoder = VoiceEncoder("cpu")
_, cont_embeds, wav_splits = encoder.embed_utterance(wav, return_partials=True, rate=16)
print(cont_embeds.shape)
clusterer = SpectralClusterer(
min_clusters=2,
max_clusters=100,
p_percentile=0.90,
gaussian_blur_sigma=1)
labels = clusterer.predict(cont_embeds)
def create_labelling(labels, wav_splits):
from resemblyzer.audio import sampling_rate
times = [((s.start + s.stop) / 2) / sampling_rate for s in wav_splits]
labelling = []
start_time = 0
for i, time in enumerate(times):
if i > 0 and labels[i] != labels[i - 1]:
temp = [str(labels[i - 1]), start_time, time]
labelling.append(tuple(temp))
start_time = time
if i == len(times) - 1:
temp = [str(labels[i]), start_time, time]
labelling.append(tuple(temp))
return labelling
labelling = create_labelling(labels, wav_splits)
这段代码帮助很大: 首先添加一个包含时间戳的time_stamps.txt文件,以修剪音频(time_stamps.txt文件应以逗号分隔)。 然后添加音频文件名和格式,它就完成了任务。我在github上找到这个,https://github.com/raotnameh/Trim_audio
相关问题 更多 >
编程相关推荐