Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Last update: Nov 25, 2022

Related tags

Deep Learning SGN

Overview

Semantic Grouping Network for Video Captioning

Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. AAAI 2021. [arxiv]

Environment

Ubuntu 16.04
CUDA 9.2
cuDNN 7.4.2
Java 8
Python 2.7.12
- PyTorch 1.1.0
- Other python packages specified in requirements.txt

Usage

1. Setup

$ pip install -r requirements.txt

2. Prepare Data

Download the GloVe Embedding from here and locate it at data/Embeddings/GloVe/GloVe_300.json.
Extract features from datasets and locate them at data/ /features/ .hdf5.

e.g. ResNet101 features of the MSVD dataset will be located at data/MSVD/features/ResNet101.hdf5.

I refer to this repo for extracting the ResNet101 features, and this repo for extracting the 3D-ResNext101 features.
Split the features into train, val, and test sets by running following commands.
```
$ python -m split.MSVD
$ python -m split.MSR-VTT
```

You can skip step 2-3 and download below files

MSVD
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]
MSR-VTT
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]

3. Prepare The Code for Evaluation

Clone the evaluation code from the official coco-evaluation repo.

$ git clone https://github.com/tylin/coco-caption.git
$ mv coco-caption/pycocoevalcap .
$ rm -rf coco-caption

4. Extract Negative Videos

$ python extract_negative_videos.py

or you can skip this step as the output files are already uploaded at data/ /metadata/neg_vids_ .json

5. Train

$ python train.py

You can change some hyperparameters by modifying config.py.

Pretrained Models - SGN(R101+RN)

*Disclaimer: The models above do not have the same weight as the models used in the paper (I trained them again because I lost).

6. Evaluate

$ python evaluate.py --ckpt_fpath

License

The source-code in this repository is released under MIT License.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Keeper for Ricochet Protocol, implemented with Apache Airflow

The implementation of the CVPR2021 paper "Structure-Aware Face Clustering on a Large-Scale Graph with 10^7 Nodes"

A Neural Net Training Interface on TensorFlow, with focus on speed + flexibility

Springer Link Download Module for Python

Code to reproduce the results in "Visually Grounded Reasoning across Languages and Cultures", EMNLP 2021.

Code release for paper: The Boombox: Visual Reconstruction from Acoustic Vibrations

A Demo server serving Bert through ONNX with GPU written in Rust with <3

A dead simple python wrapper for darknet that works with OpenCV 4.1, CUDA 10.1

Self-Learning - Books Papers, Courses & more I have to learn soon

Clinica is a software platform for clinical research studies involving patients with neurological and psychiatric diseases and the acquisition of multimodal data

TorchX is a library containing standard DSLs for authoring and running PyTorch related components for an E2E production ML pipeline.

Official Repo for Ground-aware Monocular 3D Object Detection for Autonomous Driving

Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions

Apollo optimizer in tensorflow

An open-source online reverse dictionary.

[ICCV 2021] A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation

Wav2Vec for speech recognition, classification, and audio classification

CenterFace(size of 7.3MB) is a practical anchor-free face detection and alignment method for edge devices.

Diverse Image Generation via Self-Conditioned GANs

Lightweight Python library for adding real-time object tracking to any detector.