AlphaNet Improved Training of Supernet with Alpha-Divergence

Last update: Oct 10, 2022

Related tags

Overview

AlphaNet: Improved Training of Supernet with Alpha-Divergence

This repository contains our PyTorch training code, evaluation code and pretrained models for AlphaNet.

Our implementation is largely based on AttentiveNAS. To reproduce our results, please first download the AttentiveNAS repo, and use our train_alphanet.py for training and test_alphanet.py for testing.

For more details, please see AlphaNet: Improved Training of Supernet with Alpha-Divergence by Dilin Wang, Chengyue Gong, Meng Li, Qiang Liu, Vikas Chandra.

If you find this repo useful in your research, please consider citing our work and AttentiveNAS:

@article{wang2021alphanet,
  title={AlphaNet: Improved Training of Supernet with Alpha-Divergence},
  author={Wang, Dilin and Gong, Chengyue and Li, Meng and Liu, Qiang and Chandra, Vikas},
  journal={arXiv preprint arXiv:2102.07954},
  year={2021}
}

@article{wang2020attentivenas,
  title={AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling},
  author={Wang, Dilin and Li, Meng and Gong, Chengyue and Chandra, Vikas},
  journal={arXiv preprint arXiv:2011.09011},
  year={2020}
}

Evaluation

To reproduce our results:

Please first download our pretrained AlphaNet models from a Google Drive path and put the pretrained models under your local folder ./alphanet_data

To evaluate our pre-trained AlphaNet models, from AlphaNet-A0 to A6, on ImageNet with a single GPU, please run:

python test_alphanet.py --config-file ./configs/eval_alphanet_models.yml --model a[0-6]

Expected results:

Name	MFLOPs	Top-1 (%)
AlphaNet-A0	203	77.87
AlphaNet-A1	279	78.94
AlphaNet-A2	317	79.20
AlphaNet-A3	357	79.41
AlphaNet-A4	444	80.01
AlphaNet-A5 (small)	491	80.29
AlphaNet-A5 (base)	596	80.62
AlphaNet-A6	709	80.78

Additionally, here is our pretrained supernet with KL based inplace-KD and here is our pretrained supernet without inplace-KD.

Training

To train our AlphaNet models from scratch, please run:

python train_alphanet.py --config-file configs/train_alphanet_models.yml --machine-rank ${machine_rank} --num-machines ${num_machines} --dist-url ${dist_url}

We adopt SGD training on 64 GPUs. The mini-batch size is 32 per GPU; all training hyper-parameters are specified in train_alphanet_models.yml.

Evolutionary search

In case you want to search the set of models of your own interest - we provide an example to show how to search the Pareto models for the best FLOPs vs. accuracy tradeoffs in parallel_supernet_evo_search.py; to run this example:

python parallel_supernet_evo_search.py --config-file configs/parallel_supernet_evo_search.yml

License

AlphaNet is licensed under CC-BY-NC.

Contributing

We actively welcome your pull requests! Please see CONTRIBUTING and CODE_OF_CONDUCT for more info.

AlphaNet Improved Training of Supernet with Alpha-Divergence

Related tags

Overview

AlphaNet: Improved Training of Supernet with Alpha-Divergence

Evaluation

Training

Evolutionary search

License

Contributing

Owner

Facebook Research

3.8% and 18.3% on CIFAR-10 and CIFAR-100

BERT model training impelmentation using 1024 A100 GPUs for MLPerf Training v1.1

Fully Convlutional Neural Networks for state-of-the-art time series classification

Deeper insights into graph convolutional networks for semi-supervised learning

A python implementation of Physics-informed Spline Learning for nonlinear dynamics discovery

Convnext-tf - Unofficial tensorflow keras implementation of ConvNeXt

Winning Solution in NTIRE19 Challenges on Video Restoration and Enhancement (CVPR19 Workshops) - Video Restoration with Enhanced Deformable Convolutional Networks. EDVR has been merged into BasicSR and this repo is a mirror of BasicSR.

MVGCN: a novel multi-view graph convolutional network (MVGCN) framework for link prediction in biomedical bipartite networks.

Group Fisher Pruning for Practical Network Compression(ICML2021)

From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement (CVPR'2020)

Code for the paper: Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization (https://arxiv.org/abs/2002.11798)

High level network definitions with pre-trained weights in TensorFlow

Graph Posterior Network: Bayesian Predictive Uncertainty for Node Classification (NeurIPS 2021)

LowRankModels.jl is a julia package for modeling and fitting generalized low rank models.

Must-read Papers on Physics-Informed Neural Networks.

Implementation of character based convolutional neural network

A Pose Estimator for Dense Reconstruction with the Structured Light Illumination Sensor

This project provides the code and datasets for 'CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection', CVPR 2019.

A rough implementation of the paper "A Steering Algorithm for Redirected Walking Using Reinforcement Learning"

Deep-Learning-Book-Chapter-Summaries - Attempting to make the Deep Learning Book easier to understand.