使用深度学习框架提取视频硬字幕；docker容器免安装深度学习库，使用本地api接口使得界面和后端识别分离；

Last update: Aug 06, 2022

Related tags

Overview

extract-video-subtittle

使用深度学习框架提取视频硬字幕；

本地识别无需联网；

CPU识别速度可观；

容器提供API接口；

运行环境

本项目运行环境非常好搭建，我做好了docker容器免安装各种深度学习包；

提供windows界面操作；

容器为CPU版本；

视频演示

https://www.bilibili.com/video/BV18Q4y1f774/

程序说明

1、先启动后端容器实例

docker run -d -p 6666:6666 m986883511/extract_subtitles

2、启动程序

简单介绍页面

1：点击左边按钮连接第一步启动的容器；

2：视频提取字幕的总进度

3：当前视频帧显示的位置，就是视频进度条

4：识别出来的文字会在这里显示一下

3、点击选择视频确认字幕位置

点击选择视频按钮，这时你可以拖动进度条到有字幕的位置；然后点击选择字幕区域；在视频中画一个矩形；

4、点击测试连接API

后端没问题的话，会显示已连通；此时所有步骤准备就绪

5、开始识别

点击请先完成前几步按钮，内部分为这几个步骤

本地通过ffmpeg提取视频声音保存到temp目录（0%-10%）
api通信将声音文件发送到容器内，容器内spleeter库提取声音中人声，结果保存在容器内temp目录，很耗时间，吃CPU和内存（10%-30）
api通信，将人声根据停顿分片，返回分片结果，耗较短的时间（30%-40%）
根据说话分片时间开始识别字幕（40-%100%）

当100%的时候查看temp目录就生成了和视频同名的srt字幕文件

运行后台

后端接口容器地址Docker Hub

此过程可能时间较长，您需要预先安装好好docker，并配置好docker加速器，你可能需要先docker login

docker run -d -p 6666:6666 m986883511/extract_subtitles

本项目缺少文件

因网速墙的问题，大文件推送不上去，可以参考.gitignore中写的

其他

视频提取

# 视频片段提取
ffmpeg -ss 00:15:45 -t 00:02:15 -i test/three_body_3_7.mp4 -vcodec copy -acodec copy test/3body.mp4
# 打包界面程序
C:/Python/Python38-32/Scripts/pyinstaller.exe main.spec

参考资料

本项目中深度学习源代码为/docker/backend

原作者为：https://github.com/YaoFANGUK/video-subtitle-extractor

Comments

提取人声一直没结果

视频是40多分钟的连续剧。CPU版本。之前用YaoFANGUK/video-subtitle-extractor提取字幕很成功也准确，但时间比较长。看到作者用音频分析减少了识别的帧数，所以试了一下。但在提取人声时，已经等待了近50分钟没有结果。而且CPU的占用只有1%左右，这明显不正常。用YaoFANGUK/video-subtitle-extractor整个的耗时可能都没有这么久。另外autosub也是提取音频来语音识别字幕，识别人声也很快，同样的视频几分钟就完了。麻烦作者看看是出了什么问题呢。

opened by royzengyi 2
项目咨询

Hello，我尝试了一下这个软件，感觉还是不错的，不过在实际使用中还是会有不少问题。

我是一个独立开发者，这边愿意付费或者合作来完善一下，让这个项目更具实用性，不知道你有没有兴趣呢?

没有找到联系方式，只好通过issue来试一下，你可以在看到之后删除，谢谢。

我的邮箱是yedaxia#foxmail.com

opened by YeDaxia 1

Releases(0.2.0)

0.2.0(Aug 2, 2021)

1、修复ffmepg缺少dll的问题 2、修改双击exe报错的问题，换成英文名字 3、增加config.json，可配置后端的ip和port 4、修复ffmpeg没有添加到系统PATH的bug
Source code(tar.gz)
Source code(zip)
extract-video-subtittle-v0.2.0.7z(65.84 MB)

Owner

歌者

失去人性，失去很多；失去兽性，失去一切；活着才能燃烧自己。

GitHub Repository

TUPÃ was developed to analyze electric field properties in molecular simulations

TUPÃ: Electric field analyses for molecular simulations What is TUPÃ? TUPÃ (pronounced as tu-pan) is a python algorithm that employs MDAnalysis engine

10 Jul 17, 2022

Source Code For Template-Based Named Entity Recognition Using BART

Template-Based NER Source Code For Template-Based Named Entity Recognition Using BART Training Training train.py Inference inference.py Corpus ATIS (h

174 Dec 19, 2022

PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models

Deepvoice3_pytorch PyTorch implementation of convolutional networks-based text-to-speech synthesis models: arXiv:1710.07654: Deep Voice 3: Scaling Tex

1.8k Jan 08, 2023

Federated learning on graph, especially on graph neural networks (GNNs), knowledge graph, and private GNN.

198 Dec 20, 2022

Source Code for DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank Utterances (https://arxiv.org/pdf/2012.01775.pdf)

DialogBERT This is a PyTorch implementation of the DialogBERT model described in DialogBERT: Neural Response Generation via Hierarchical BERT with Dis

67 Jan 06, 2023

An open-source, low-cost, image-based weed detection device for fallow scenarios.

Welcome to the OpenWeedLocator (OWL) project, an opensource hardware and software green-on-brown weed detector that uses entirely off-the-shelf compon

145 Jan 05, 2023

Accelerated deep learning R&D

Accelerated deep learning R&D PyTorch framework for Deep Learning research and development. It focuses on reproducibility, rapid experimentation, and

3.1k Jan 06, 2023

Easy to use Audio Tagging in PyTorch

Audio Classification, Tagging & Sound Event Detection in PyTorch Progress: Fine-tune on audio classification Fine-tune on audio tagging Fine-tune on s

15 Dec 22, 2022

🤖 A Python library for learning and evaluating knowledge graph embeddings

PyKEEN PyKEEN (Python KnowlEdge EmbeddiNgs) is a Python package designed to train and evaluate knowledge graph embedding models (incorporating multi-m

1.1k Jan 09, 2023

Norm-based Analysis of Transformer

Norm-based Analysis of Transformer Implementations for 2 papers introducing to analyze Transformers using vector norms: Kobayashi+'20 Attention is Not

52 Dec 05, 2022

This is a student data management application developed in Python and TKinter. It utilizes the TKinter pillow library to include images to buttons. I've separated TKinter elements into their own individual classes. The user can change the smilely face color for each button individually or by entire row.

Smiley Face Cube Display Table of Contents Project Description Getting Started Prerequisites Installation & Deployment Additional Documentation Projec

0 Aug 04, 2021

使用深度学习框架提取视频硬字幕；docker容器免安装深度学习库，使用本地api接口使得界面和后端识别分离；

Related tags

Overview

extract-video-subtittle

运行环境

视频演示

程序说明

运行后台

本项目缺少文件

其他

参考资料

You might also like...

Comments

提取人声一直没结果

项目咨询

Releases(0.2.0)

0.2.0(Aug 2, 2021)

Owner

歌者

TUPÃ was developed to analyze electric field properties in molecular simulations

Source Code For Template-Based Named Entity Recognition Using BART

PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models

Federated learning on graph, especially on graph neural networks (GNNs), knowledge graph, and private GNN.

Source Code for DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank Utterances (https://arxiv.org/pdf/2012.01775.pdf)

An open-source, low-cost, image-based weed detection device for fallow scenarios.

Accelerated deep learning R&D

Easy to use Audio Tagging in PyTorch

🤖 A Python library for learning and evaluating knowledge graph embeddings

Norm-based Analysis of Transformer

Understanding and Overcoming the Challenges of Efficient Transformer Quantization

A framework for annotating 3D meshes using the predictions of a 2D semantic segmentation model.

Code for “ACE-HGNN: Adaptive Curvature ExplorationHyperbolic Graph Neural Network”

This is the repo of the manuscript "Dual-branch Attention-In-Attention Transformer for speech enhancement"

[ArXiv 2021] Data-Efficient Instance Generation from Instance Discrimination

Manipulation OpenAI Gym environments to simulate robots at the STARS lab

[ICCV 2021] Released code for Causal Attention for Unbiased Visual Recognition

Semi-supervised Semantic Segmentation with Directional Context-aware Consistency (CVPR 2021)

Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising