ML-Decoder: Scalable and Versatile Classification Head

Last update: Jan 04, 2023

Related tags

Deep Learning ML_Decoder

Overview

ML-Decoder: Scalable and Versatile Classification Head

Paper

Official PyTorch Implementation

Tal Ridnik, Gilad Sharir, Avi Ben-Cohen, Emanuel Ben-Baruch, Asaf Noy
DAMO Academy, Alibaba Group

Abstract

In this paper, we introduce ML-Decoder, a new attention-based classification head. ML-Decoder predicts the existence of class labels via queries, and enables better utilization of spatial data compared to global average pooling. By redesigning the decoder architecture, and using a novel group-decoding scheme, ML-Decoder is highly efficient, and can scale well to thousands of classes. Compared to using a larger backbone, ML-Decoder consistently provides a better speed-accuracy trade-off. ML-Decoder is also versatile - it can be used as a drop-in replacement for various classification heads, and generalize to unseen classes when operated with word queries. Novel query augmentations further improve its generalization ability. Using ML-Decoder, we achieve state-of-the-art results on several classification tasks: on MS-COCO multi-label, we reach 91.4% mAP; on NUS-WIDE zero-shot, we reach 31.1% ZSL mAP; and on ImageNet single-label, we reach with vanilla ResNet50 backbone a new top score of 80.7%, without extra data or distillation.

ML-Decoder Implementation

ML-Decoder implementation is available here. It can be easily integrated into any backbone using this example code:

ml_decoder_head = MLDecoder(num_classes) # initilization

spatial_embeddings = self.backbone(input_image) # backbone generates spatial embeddings      
 
logits = ml_decoder_head(spatial_embeddings) # transfrom spatial embeddings to logits

Training Code

We will share a full reproduction code for the article results.

Multi-label Training Code

A reproduction code for MS-COCO multi-label:

python train.py  \
--data=/home/datasets/coco2014/ \
--model_name=tresnet_l \
--image_size=448

Single-label Training Code

Our single-label training code uses the excellent timm repo. Reproduction code is currently from a fork, we will work toward a full merge to the main repo.

git clone https://github.com/mrT23/pytorch-image-models.git

This is the code for A2 configuration training, with ML-Decoder (--use-ml-decoder-head=1):

python -u -m torch.distributed.launch --nproc_per_node=8 \
--nnodes=1 \
--node_rank=0 \
./train.py \
/data/imagenet/ \
--amp \
-b=256 \
--epochs=300 \
--drop-path=0.05 \
--opt=lamb \
--weight-decay=0.02 \
--sched='cosine' \
--lr=4e-3 \
--warmup-epochs=5 \
--model=resnet50 \
--aa=rand-m7-mstd0.5-inc1 \
--reprob=0.0 \
--remode='pixel' \
--mixup=0.1 \
--cutmix=1.0 \
--aug-repeats 3 \
--bce-target-thresh 0.2 \
--smoothing=0 \
--bce-loss \
--train-interpolation=bicubic \
--use-ml-decoder-head=1

ZSL Training Code

Reproduction code for ZSL is WIP.

Citation

@misc{ridnik2021mldecoder,
      title={ML-Decoder: Scalable and Versatile Classification Head}, 
      author={Tal Ridnik and Gilad Sharir and Avi Ben-Cohen and Emanuel Ben-Baruch and Asaf Noy},
      year={2021},
      eprint={2111.12933},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

ML-Decoder: Scalable and Versatile Classification Head

Related tags

Overview

ML-Decoder: Scalable and Versatile Classification Head

ML-Decoder Implementation

Training Code

Multi-label Training Code

Single-label Training Code

ZSL Training Code

Citation

Owner

Recursive Bayesian Networks

[NeurIPS'21] "AugMax: Adversarial Composition of Random Augmentations for Robust Training" by Haotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu, Animashree Anandkumar, and Zhangyang Wang.

Meta Representation Transformation for Low-resource Cross-lingual Learning

FuseDream: Training-Free Text-to-Image Generationwith Improved CLIP+GAN Space OptimizationFuseDream: Training-Free Text-to-Image Generationwith Improved CLIP+GAN Space Optimization

Awesome Artificial Intelligence, Machine Learning and Deep Learning as we learn it

Kaggle | 9th place single model solution for TGS Salt Identification Challenge

A PaddlePaddle version of Neural Renderer, refer to its PyTorch version

CellRank's reproducibility repository.

Self-Supervised depth kalilia

Image super-resolution through deep learning

Task-related Saliency Network For Few-shot learning

Implementation supporting the ICCV 2017 paper "GANs for Biological Image Synthesis"

Use Python, OpenCV, and MediaPipe to control a keyboard with facial gestures

🚀 An end-to-end ML applications using PyTorch, W&B, FastAPI, Docker, Streamlit and Heroku

Curvlearn, a Tensorflow based non-Euclidean deep learning framework.

In this work, we will implement some basic but important algorithm of machine learning step by step.

Open source repository for the code accompanying the paper 'PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations'.

CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient Estimator

An implementation of a discriminant function over a normal distribution to help classify datasets.

AI创造营：Metaverse启动机之重构现世，结合PaddlePaddle 和 Wechaty 创造自己的聊天机器人

ML-Decoder: Scalable and Versatile Classification Head

Related tags

Overview

ML-Decoder: Scalable and Versatile Classification Head

ML-Decoder Implementation

Training Code

Multi-label Training Code

Single-label Training Code

ZSL Training Code

Citation

Owner

Recursive Bayesian Networks

[NeurIPS'21] "AugMax: Adversarial Composition of Random Augmentations for Robust Training" by Haotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu, Animashree Anandkumar, and Zhangyang Wang.

Meta Representation Transformation for Low-resource Cross-lingual Learning

FuseDream: Training-Free Text-to-Image Generationwith Improved CLIP+GAN Space OptimizationFuseDream: Training-Free Text-to-Image Generationwith Improved CLIP+GAN Space Optimization

Awesome Artificial Intelligence, Machine Learning and Deep Learning as we learn it

Kaggle | 9th place single model solution for TGS Salt Identification Challenge

A PaddlePaddle version of Neural Renderer, refer to its PyTorch version

CellRank's reproducibility repository.

Self-Supervised depth kalilia

Image super-resolution through deep learning

Task-related Saliency Network For Few-shot learning

Implementation supporting the ICCV 2017 paper "GANs for Biological Image Synthesis"

Use Python, OpenCV, and MediaPipe to control a keyboard with facial gestures

🚀 An end-to-end ML applications using PyTorch, W&B, FastAPI, Docker, Streamlit and Heroku

Curvlearn, a Tensorflow based non-Euclidean deep learning framework.

In this work, we will implement some basic but important algorithm of machine learning step by step.

Open source repository for the code accompanying the paper 'PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations'.

CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient Estimator

An implementation of a discriminant function over a normal distribution to help classify datasets.

AI创造营 ：Metaverse启动机之重构现世，结合PaddlePaddle 和 Wechaty 创造自己的聊天机器人

AI创造营：Metaverse启动机之重构现世，结合PaddlePaddle 和 Wechaty 创造自己的聊天机器人