A Pytorch implementation of MoveNet from Google. Include training code and pre-train model.

Last update: Dec 26, 2022

Related tags

Overview

Movenet.Pytorch

Intro

MoveNet is an ultra fast and accurate model that detects 17 keypoints of a body. This is A Pytorch implementation of MoveNet from Google. Include training code and pre-train model.

Google just release pre-train models(tfjs or tflite), which cannot be converted to some CPU inference framework such as NCNN,Tengine,MNN,TNN, and we can not add our own custom data to finetune, so there is this repo.

How To Run

1.Download COCO dataset2017 from https://cocodataset.org/. (You need train2017.zip, val2017.zip and annotations.)Unzip to movenet.pytorch/data/ like this:

├── data
    ├── annotations (person_keypoints_train2017.json, person_keypoints_val2017.json, ...)
    ├── train2017   (xx.jpg, xx.jpg,...)
    └── val2017     (xx.jpg, xx.jpg,...)

2.Make data to our data format.

python scripts/make_coco_data_17keypooints.py

Our data format: JSON file
Keypoints order:['nose', 'left_eye', 'right_eye', 'left_ear', 'right_ear', 
    'left_shoulder', 'right_shoulder', 'left_elbow', 'right_elbow', 'left_wrist', 
    'right_wrist', 'left_hip', 'right_hip', 'left_knee', 'right_knee', 'left_ankle', 
    'right_ankle']

One item:
[{"img_name": "0.jpg",
  "keypoints": [x0,y0,z0,x1,y1,z1,...],
  #z: 0 for no label, 1 for labeled but invisible, 2 for labeled and visible
  "center": [x,y],
  "bbox":[x0,y0,x1,y1],
  "other_centers": [[x0,y0],[x1,y1],...],
  "other_keypoints": [[[x0,y0],[x1,y1],...],[[x0,y0],[x1,y1],...],...], #lenth = num_keypoints
 },
 ...
]

3.You can add your own data to the same format.

4.After putting data at right place, you can start training

python train.py

5.After training finished, you need to change the test model path to test. Such as this in predict.py

run_task.modelLoad("output/xxx.pth")

6.run predict to show predict result, or run evaluate.py to compute my acc on test dataset.

python predict.py

7.Convert to onnx.

python pth2onnx.py

Training Results

Some good samples

Some bad cases

Tips to improve

1. Focus on data

Add COCO2014. (But as I know it has some duplicate data of COCO2017, and I don't know if google use it.)
Clean the croped COCO2017 data. (Some img just have little points, such as big face, big body,etc.MoveNet is a small network, COCO data is a little hard for it.)
Add some yoga, fitness, and dance videos frame from YouTube. (Highly Recommened! Cause Google did this on their Movenet and said 'Evaluations on the Active validation dataset show a significant performance boost relative to identical architectures trained using only COCO. ')

2. Change backbone

Try to ransfer Mobilenetv2(original Movenet) to Mobilenetv3 or Shufflenetv2 may get a litte improvement.If you just wanna reproduce the original Movenet, u can ignore this.

3. More fancy loss

Surely this is a muti-task learning. So add some loss to learn together may improve the performence. (Such as BoneLoss which I have added.) And we can never know how Google trained, cause we cannot see it from the pre-train tflite model file, so you can try any loss function you like.

4. Data Again

I just wanna you know the importance of the data. The more time you spend on clean data and add new data, the better performance your model will get! (While tips 2 and 3 may not.)

A Pytorch implementation of MoveNet from Google. Include training code and pre-train model.

Related tags

Overview

Movenet.Pytorch

Intro

How To Run

Training Results

Some good samples

Some bad cases

Tips to improve

1. Focus on data

2. Change backbone

3. More fancy loss

4. Data Again

Resource

Owner

Mr.Fire

Python Algorithm Interview Book Review

Official Pytorch implementation for 2021 ICCV paper "Learning Motion Priors for 4D Human Body Capture in 3D Scenes" and trained models / data

StyleSwin: Transformer-based GAN for High-resolution Image Generation

YOLO-v5 기반 단안 카메라의 영상을 활용해 차간 거리를 일정하게 유지하며 주행하는 Adaptive Cruise Control 기능 구현

Pytorch version of VidLanKD: Improving Language Understanding viaVideo-Distilled Knowledge Transfer

Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System

QSYM: A Practical Concolic Execution Engine Tailored for Hybrid Fuzzing

My 1st place solution at Kaggle Hotel-ID 2021

Multi-Output Gaussian Process Toolkit

LSTM model trained on a small dataset of 3000 names written in PyTorch

This repository contains a Ruby API for utilizing TensorFlow.

(NeurIPS 2021) Pytorch implementation of paper "Re-ranking for image retrieval and transductive few-shot classification"

[ICCV21] Self-Calibrating Neural Radiance Fields

All course materials for the Zero to Mastery Machine Learning and Data Science course.

Cross-media Structured Common Space for Multimedia Event Extraction (ACL2020)

Official Pytorch implementation of the paper "MotionCLIP: Exposing Human Motion Generation to CLIP Space"

Joint-task Self-supervised Learning for Temporal Correspondence (NeurIPS 2019)

Back to Event Basics: SSL of Image Reconstruction for Event Cameras

A general python framework for visual object tracking and video object segmentation, based on PyTorch

Official code of CVPR 2021's PLOP: Learning without Forgetting for Continual Semantic Segmentation