Official Pytorch implementation for video neural representation (NeRV)

Last update: Dec 28, 2022

Related tags

Overview

NeRV: Neural Representations for Videos (NeurIPS 2021)

Project Page | Paper | UVG Data

Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser-Nam Lim, Abhinav Shrivastava
This is the official implementation of the paper "NeRV: Neural Representations for Videos ".

Get started

We run with Python 3.8, you can set up a conda environment with all dependencies like so:

pip install -r requirements.txt

High-Level structure

The code is organized as follows:

train_nerv.py includes a generic traiing routine.
model_nerv.py contains the dataloader and neural network architecure
data/ directory video/imae dataset, we provide big buck bunny here
checkpoint/ directory contains some pre-trained model on big buck bunny dataset
log files (tensorboard, txt, state_dict etc.) will be saved in output directory (specified by --outf)

Reproducing experiments

Training experiments

The NeRV-S experiment on 'big buck bunny' can be reproduced with

python train_nerv.py -e 300 --cycles 1  --lower-width 96 --num-blocks 1 --dataset bunny --frame_gap 1 \
    --outf bunny_ab --embed 1.25_40 --stem_dim_num 512_1  --reduction 2  --fc_hw_dim 9_16_26 --expansion 1  \
    --single_res --loss Fusion6   --warmup 0.2 --lr_type cosine  --strides 5 2 2 2 2  --conv_type conv \
    -b 1  --lr 0.0005 --norm none --act swish

Evaluation experiments

To evaluate pre-trained model, just add --eval_Only and specify model path with --weight, you can specify model quantization with --quant_bit [bit_lenght], yuo can test decoding speed with --eval_fps, below we preovide sample commends for NeRV-S on bunny dataset

python train_nerv.py -e 300 --cycles 1  --lower-width 96 --num-blocks 1 --dataset bunny --frame_gap 1 \
    --outf bunny_ab --embed 1.25_40 --stem_dim_num 512_1  --reduction 2  --fc_hw_dim 9_16_26 --expansion 1  \
    --single_res --loss Fusion6   --warmup 0.2 --lr_type cosine  --strides 5 2 2 2 2  --conv_type conv \
    -b 1  --lr 0.0005 --norm none  --act swish \
    --weight checkpoints/nerv_S.pth --eval_only

Dump predictions with pre-trained model

To evaluate pre-trained model, just add --eval_Only and specify model path with --weight

python train_nerv.py -e 300 --cycles 1  --lower-width 96 --num-blocks 1 --dataset bunny --frame_gap 1 \
    --outf bunny_ab --embed 1.25_40 --stem_dim_num 512_1  --reduction 2  --fc_hw_dim 9_16_26 --expansion 1  \
    --single_res --loss Fusion6   --warmup 0.2 --lr_type cosine  --strides 5 2 2 2 2  --conv_type conv \
    -b 1  --lr 0.0005 --norm none  --act swish \
   --weight checkpoints/nerv_S.pth --eval_only  --dump_images

Citation

If you find our work useful in your research, please cite:

@inproceedings{hao2021nerv,
    author = {Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser-Nam Lim, Abhinav Shrivastava },
    title = {NeRV: Neural Representations for Videos s},
    booktitle = {NeurIPS},
    year={2021}
}

Contact

If you have any questions, please feel free to email the authors.

Official Pytorch implementation for video neural representation (NeRV)

Related tags

Overview

NeRV: Neural Representations for Videos (NeurIPS 2021)

Project Page | Paper | UVG Data

Get started

High-Level structure

Reproducing experiments

Training experiments

Evaluation experiments

Dump predictions with pre-trained model

Citation

Contact

Owner

hao

AITom is an open-source platform for AI driven cellular electron cryo-tomography analysis.

Official implementation of the paper ``Unifying Nonlocal Blocks for Neural Networks'' (ICCV'21)

Sharpened cosine similarity torch - A Sharpened Cosine Similarity layer for PyTorch

Real-Time-Student-Attendence-System - Real Time Student Attendence System

ML-based medical imaging using Azure

A Comparative Framework for Multimodal Recommender Systems

UFPR-ADMR-v2 Dataset

Java and SHACL code commented in the paper "Towards compliance checking in reified I/O logic via SHACL" submitted to ICAIL 2021

[NeurIPS2021] Code Release of K-Net: Towards Unified Image Segmentation

Unrolled Variational Bayesian Algorithm for Image Blind Deconvolution

Spatial Action Maps for Mobile Manipulation (RSS 2020)

Code for the paper titled "Prabhupadavani: A Code-mixed Speech Translation Data for 25 languages"

PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models

Code of the paper "Shaping Visual Representations with Attributes for Few-Shot Learning (ASL)".

RepVGG: Making VGG-style ConvNets Great Again

Mosaic of Object-centric Images as Scene-centric Images (MosaicOS) for long-tailed object detection and instance segmentation.

Repository for the paper "PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation", CVPR 2021.

This is the dataset and code release of the OpenRooms Dataset.

基于Paddle框架的fcanet复现

PyTorch implementation of Off-policy Learning in Two-stage Recommender Systems