VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Last update: Dec 12, 2022

Related tags

Overview

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Usage

First, install PyTorch 1.7.1+, torchvision 0.8.2+ and other required packages as follows：

conda install -c pytorch pytorch torchvision
pip install timm==0.3.2
pip install ftfy regex tqdm
pip install git+https://github.com/openai/CLIP.git
pip install mmcv==1.3.14

Data preparation

ImageNet-LT

Download and extract ImageNet train and val images from here. The directory structure is the standard layout for the torchvision datasets.ImageFolder, and the training and validation data is expected to be in the train/ folder and val/ folder respectively.

Then download and extract the wiki text into the same directory, and the directory tree of data is expected to be like this:

./data/imagenet/
  train/
    class1/
      img1.jpeg
    class2/
      img2.jpeg
  val/
    class1/
      img3.jpeg
    class2/
      img4.jpeg
  wiki/
  	desc_1.txt
  ImageNet_LT_test.txt
  ImageNet_LT_train.txt
  ImageNet_LT_val.txt
  labels.txt

After that, download the CLIP's pretrained weight RN50.pt and ViT-B-16.pt into the pretrained directory from https://github.com/openai/CLIP.

Places-LT

Download the places365_standard data from here.

Then download and extract the wiki text into the same directory. The directory tree of data is expected to be like this (almost the same as ImageNet-LT):

./data/places/
  train/
    class1/
      img1.jpeg
    class2/
      img2.jpeg
  val/
    class1/
      img3.jpeg
    class2/
      img4.jpeg
  wiki/
  	desc_1.txt
  Places_LT_test.txt
  Places_LT_train.txt
  Places_LT_val.txt
  labels.txt

iNaturalist 2018

Download the iNaturalist 2018 data from here.

Then download and extract the wiki text into the same directory. The directory tree of data is expected to be like this:

./data/iNat/
  train_val2018/
  wiki/
  	desc_1.txt
  categories.json
  test2018.json
  train2018.json
  val.json

Evaluation

To evaluate VL-LTR with a single GPU run:

Pre-training stage

bash eval.sh ${CONFIG_PATH} 1 --eval-pretrain

Fine-tuning stage:

bash eval.sh ${CONFIG_PATH} 1

The ${CONFIG_PATH} is the relative path of the corresponding configuration file in the config directory.

Training

To train VL-LTR on a single node with 8 GPUs for:

Pre-training stage, run:

bash dist_train_arun.sh ${PARTITION} ${CONFIG_PATH} 8

Fine-tuning stage:
- First, calculate the $\mathcal L_{\text{lin}}$ of each sentence for AnSS method by running this:
```
bash eval.sh ${CONFIG_PATH} 1 --eval-pretrain --select
```
- then, running this:
```
bash dist_train_arun.sh ${PARTITION} ${CONFIG_PATH} 8
```

The ${CONFIG_PATH} is the relative path of the corresponding configuration file in the config directory.

Results

Below list our model's performance on ImageNet-LT, Places-LT, and iNaturalist 2018.

Dataset	Backbone	Top-1 Accuracy	Download
ImageNet-LT	ResNet-50	70.1	Weights
ImageNet-LT	ViT-Base-16	77.2	Weights
Places-LT	ResNet-50	48.0	Weights
Places-LT	ViT-Base-16	50.1	Weights
iNaturalist 2018	ResNet-50	74.6	Weights
iNaturalist 2018	ViT-Base-16	76.8	Weights

For more detailed information, please refer to our paper directly.

Citation

If you are interested in our work, please cite as follows:

@article{tian2021vl,
  title={VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition},
  author={Tian, Changyao and Wang, Wenhai and Zhu, Xizhou and Wang, Xiaogang and Dai, Jifeng and Qiao, Yu},
  journal={arXiv preprint arXiv:2111.13579},
  year={2021}
}

License

This repository is released under the Apache 2.0 license as found in the LICENSE file.

Code for the AAAI-2022 paper: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification (AAAI 2022) Prerequisite PyTorch = 1.2.0 P

16 Dec 14, 2022

Pytorch implementation of the AAAI 2022 paper "Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification"

[AAAI22] Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification We point out the overlooked unbiasedness in long-tailed clas

28 Oct 18, 2022

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks We provide the code (in PyTorch) and datasets for our paper "On Size-Orient

4 Jun 18, 2022

Speech Emotion Recognition with Fusion of Acoustic- and Linguistic-Feature-Based Decisions

APSIPA-SER-with-A-and-T This code is the implementation of Speech Emotion Recognition (SER) with acoustic and linguistic features. The network model i

3 Jan 4, 2023

A weakly-supervised scene graph generation codebase. The implementation of our CVPR2021 paper ``Linguistic Structures as Weak Supervision for Visual Scene Graph Generation''

README.md shall be finished soon. WSSGG 0 Overview 1 Installation 1.1 Faster-RCNN 1.2 Language Parser 1.3 GloVe Embeddings 2 Settings 2.1 VG-GT-Graph

35 Nov 20, 2022

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

105 Nov 7, 2022

Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"

ResDAVEnet-VQ Official PyTorch implementation of Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech What is in this repo? M

21 Aug 23, 2022

[ICCV2021] Official code for "Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition"

CTR-GCN This repo is the official implementation for Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition. The pap

148 Dec 16, 2022

Official implementation for CVPR 2021 paper: Adaptive Class Suppression Loss for Long-Tail Object Detection

Adaptive Class Suppression Loss for Long-Tail Object Detection This repo is the official implementation for CVPR 2021 paper: Adaptive Class Suppressio

67 Dec 4, 2022

Comments

Problem about running eval.sh

""" #!/usr/bin/env bash set -x

export NCCL_LL_THRESHOLD=0

CONFIG=$1 GPUS=$1 CPUS=$[GPUS*2] PORT=${PORT:-8886}

CONFIG_NAME=${CONFIG##/} CONFIG_NAME=${CONFIG_NAME%.}

OUTPUT_DIR="./checkpoints/eval" if [ ! -d $OUTPUT_DIR ]; then mkdir ${OUTPUT_DIR} fi

python -u main.py
--port=$PORT
--num_workers 4
--resume "./checkpoints/${CONFIG_NAME}/checkpoint.pth"
--output-dir ${OUTPUT_DIR}
--config $CONFIG ${@:3}
--eval
2>&1 | tee -a ${OUTPUT_DIR}/train.log """ I have two A100, so set GPUS is 2. All other settings according to ReadME.md but I got a problem when running eval.sh """ File "eval.sh", line 4 export NCCL_LL_THRESHOLD=0 ^ SyntaxError: invalid syntax

"""

opened by euminds 2
Mismatch between code and diagram in paper for the fine-tuning phase

In fig 3, stage 2 from the paper, it looks like value for the attention is calculated based on Vision and language (Q is vision, K is language) and then applied to the language (V). But in the code, the attention is applied to the visual features. Can you verify which one is the correct way? @ChangyaoTian

opened by rahulvigneswaran 0

pre-trained weights with TorchScript?

Hello, Thanks for the great work! May I ask if it's possible for you to also provide the checkpoint weight in a TorchScript version?

It's something like:

import torch
import torchvision.models as models

model = models.resnet50()
traced = torch.jit.trace(model, (torch.rand(4, 3, 224, 224),))
torch.jit.save(traced, "test.pt")

# load model
model = torch.jit.load("test.pt")

opened by xinleihe 0

Releases(ECCV-2022-video)

ECCV-2022-video(Oct 1, 2022)

Source code(tar.gz)
Source code(zip)
4849.mp4(15.24 MB)
ECCV-2022(Oct 1, 2022)

Source code(tar.gz)
Source code(zip)
4849.pdf(2.66 MB)
text-corpus(Jan 12, 2022)

The text corpus for ImageNet-LT, Places-LT, and iNaturalist2018, including both wiki data and the class labels.
Source code(tar.gz)
Source code(zip)
imagenet.zip(7.29 MB)
iNat.zip(13.12 MB)
places.zip(2.67 MB)
checkpoints(Jan 12, 2022)

model weights for 3 datasets, including both 2 stages, using vit16 and resnet50 backbones respectively.
Source code(tar.gz)
Source code(zip)
imageNet-LT_r50.zip(1382.52 MB)
imageNet-LT_vit16.zip(1920.94 MB)
inat_finetune_r50.zip(1754.00 MB)
inat_finetune_vit16.zip(1566.75 MB)
inat_pretrain_r50.zip(1008.10 MB)
inat_pretrain_vit16.zip(1512.80 MB)
places_r50.zip(1252.33 MB)
places_vit16.zip(1843.92 MB)

Owner

GitHub Repository

Official Implementation for Fast Training of Neural Lumigraph Representations using Meta Learning.

Fast Training of Neural Lumigraph Representations using Meta Learning Project Page | Paper | Data Alexander W. Bergman, Petr Kellnhofer, Gordon Wetzst

39 Oct 08, 2022

Jarvis Project is a basic virtual assistant that uses TensorFlow for learning.

Jarvis_proyect Jarvis Project is a basic virtual assistant that uses TensorFlow for learning. Latest version 0.1 Features: Good morning protocol Tell

3 Aug 31, 2022

Exploring whether attention is necessary for vision transformers

Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet Paper/Report TL;DR We replace the attention layer in a v

461 Jan 07, 2023

It's a implement of this paper：Relation extraction via Multi-Level attention CNNs

Relation Classification via Multi-Level Attention CNNs It's a implement of this paper：Relation Classification via Multi-Level Attention CNNs. Training

2 Nov 04, 2022

Source code for 2021 ICCV paper "In-the-Wild Single Camera 3D Reconstruction Through Moving Water Surfaces"

In-the-Wild Single Camera 3D Reconstruction Through Moving Water Surfaces This is the PyTorch implementation for 2021 ICCV paper "In-the-Wild Single C

27 Dec 06, 2022

Code to reproduce results from the paper "AmbientGAN: Generative models from lossy measurements"

AmbientGAN: Generative models from lossy measurements This repository provides code to reproduce results from the paper AmbientGAN: Generative models

87 Oct 19, 2022

A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation (ICCV 2021)

A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation (ICCV 2021) This repository contains the official implemen

81 Dec 14, 2022

Offcial implementation of "A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction, ICCV-2021".

HF2-VAD Offcial implementation of "A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Predictio

76 Dec 21, 2022

Efficient electromagnetic solver based on rigorous coupled-wave analysis for 3D and 2D multi-layered structures with in-plane periodicity

Efficient electromagnetic solver based on rigorous coupled-wave analysis for 3D and 2D multi-layered structures with in-plane periodicity, such as gratings, photonic-crystal slabs, metasurfaces, surf

17 Dec 19, 2022

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Related tags

Overview

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Usage

Data preparation

ImageNet-LT

Places-LT

iNaturalist 2018

Evaluation

Training

Results

Citation

License

You might also like...

Code for the AAAI-2022 paper: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

Pytorch implementation of the AAAI 2022 paper "Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification"

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks

Speech Emotion Recognition with Fusion of Acoustic- and Linguistic-Feature-Based Decisions

A weakly-supervised scene graph generation codebase. The implementation of our CVPR2021 paper ``Linguistic Structures as Weak Supervision for Visual Scene Graph Generation''

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"

[ICCV2021] Official code for "Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition"

Official implementation for CVPR 2021 paper: Adaptive Class Suppression Loss for Long-Tail Object Detection

Comments

Problem about running eval.sh

Mismatch between code and diagram in paper for the fine-tuning phase

pre-trained weights with TorchScript?

Releases(ECCV-2022-video)

ECCV-2022-video(Oct 1, 2022)

ECCV-2022(Oct 1, 2022)

text-corpus(Jan 12, 2022)

checkpoints(Jan 12, 2022)

Owner

Official Implementation for Fast Training of Neural Lumigraph Representations using Meta Learning.

Jarvis Project is a basic virtual assistant that uses TensorFlow for learning.

Exploring whether attention is necessary for vision transformers

It's a implement of this paper：Relation extraction via Multi-Level attention CNNs

Source code for 2021 ICCV paper "In-the-Wild Single Camera 3D Reconstruction Through Moving Water Surfaces"

Code to reproduce results from the paper "AmbientGAN: Generative models from lossy measurements"

A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation (ICCV 2021)

Offcial implementation of "A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction, ICCV-2021".

Efficient electromagnetic solver based on rigorous coupled-wave analysis for 3D and 2D multi-layered structures with in-plane periodicity

Multi-scale discriminator feature-wise loss function

This is the repo for Uncertainty Quantification 360 Toolkit.

Nest - A flexible tool for building and sharing deep learning modules

交互式标注软件，暂定名 iann

Madanalysis5 - A package for event file analysis and recasting of LHC results

Source-to-Source Debuggable Derivatives in Pure Python

Eff video representation - Efficient video representation through neural fields

Software Platform for solving and manipulating multiparametric programs in Python

The official implementation of paper Siamese Transformer Pyramid Networks for Real-Time UAV Tracking, accepted by WACV22

Official Implementation of "Learning Disentangled Behavior Embeddings"

Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.