Aggragrating Nested Transformer Official Jax Implementation

Last update: Dec 20, 2022

Overview

Aggragrating Nested Transformer Official Jax Implementation

NesT is a simple method, which aggragrates nested local transformers on image blocks. The idea makes vision transformers attain better accuracy, data efficiency, and convergence on the ImageNet benchmark. NesT can be scaled to small datasets to match convnet accuracy.

This is not an officially supported Google product.

Pretrained Models and Results

Model	Accuracy	Checkpoint path
Nest-B	83.8	gs://gresearch/nest-checkpoints/nest-b_imagenet
Nest-S	83.3	gs://gresearch/nest-checkpoints/nest-s_imagenet
Nest-T	81.5	gs://gresearch/nest-checkpoints/nest-t_imagenet

Note: Accuracy is evaluated on the ImageNet2012 validation set.

Tensorbord.dev

See ImageNet training logs at Tensorboard.dev.

Colab

Colab is available for test: https://colab.sandbox.google.com/github/google-research/nested-transformer/blob/main/colab.ipynb

Instruction on Image Classification

Environment setup

virtualenv -p python3 --system-site-packages nestenv
source nestenv/bin/activate

pip install -r requirements.txt

Evaluate on ImageNet

At the first time, download ImageNet following tensorflow_datasets instruction from command lines. Optionally, download all pre-trained checkpoints

bash ./checkpoints/download_checkpoints.sh

Run the evaluation script to evaluate NesT-B.

python main.py --config configs/imagenet_nest.py --config.eval_only=True \
  --config.init_checkpoint="./checkpoints/nest-b_imagenet/ckpt.39" \
  --workdir="./checkpoints/nest-t_imagenet_eval"

Train on ImageNet

The default configuration trains NesT-B on TPUv2 8x8 with per device batch size 16.

python main.py --config configs/imagenet_nest.py --jax_backend_target=<TPU_IP_ADDRESS> --jax_xla_backend="tpu_driver" --workdir="./checkpoints/nest-b_imagenet"

Note: See jax/cloud_tpu_colab for info about TPU_IP_ADDRESS.

Train NesT-T on 8 GPUs.

python main.py --config configs/imagenet_nest_tiny.py --workdir="./checkpoints/nest-t_imagenet_8gpu"

The codebase does not support multi-node GPU training (>8 GPUs). The models reported in our paper is trained using TPU with 1024 total batch size.

Train on CIFAR

# Recommend to train on 2 GPUs. Training NesT-T can use 1 GPU.
CUDA_VISIBLE_DEVICES=0,1 python  main.py --config configs/cifar_nest.py --workdir="./checkpoints/nest_cifar"

Cite

@inproceedings{zhang2021aggregating,
  title={Aggregating Nested Transformers},
  author={Zizhao Zhang and Han Zhang and Long Zhao and Ting Chen and Tomas Pfister},
  booktitle={arXiv preprint arXiv:2105.12723},
  year={2021}
}

Aggragrating Nested Transformer Official Jax Implementation

Related tags

Overview

Aggragrating Nested Transformer Official Jax Implementation

Pretrained Models and Results

Tensorbord.dev

Colab

Instruction on Image Classification

Environment setup

Evaluate on ImageNet

Train on ImageNet

Train NesT-T on 8 GPUs.

Train on CIFAR

Cite

Owner

Google Research

Next-Best-View Estimation based on Deep Reinforcement Learning for Active Object Classification

Very large and sparse networks appear often in the wild and present unique algorithmic opportunities and challenges for the practitioner

Neural style in TensorFlow! 🎨

Collection of in-progress libraries for entity neural networks.

Underwater industrial application yolov5m6

The original implementation of TNDM used in the NeurIPS 2021 paper (no longer being updated)

Manage the availability of workspaces within Frappe/ ERPNext (sidebar) based on user-roles

Official code for On Path Integration of Grid Cells: Group Representation and Isotropic Scaling (NeurIPS 2021)

Learning from Guided Play: A Scheduled Hierarchical Approach for Improving Exploration in Adversarial Imitation Learning Source Code

PyTorch implementation of a Real-ESRGAN model trained on custom dataset

Hunt down social media accounts by username across social networks

ICLR 2021, Fair Mixup: Fairness via Interpolation

Active window border replacement for window managers.

Scalable Optical Flow-based Image Montaging and Alignment

Multi-atlas segmentation (MAS) is a promising framework for medical image segmentation

PyTorch code for the paper: FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning

Implementation for the paper 'YOLO-ReT: Towards High Accuracy Real-time Object Detection on Edge GPUs'

An efficient 3D semantic segmentation framework for Urban-scale point clouds like SensatUrban, Campus3D, etc.

Pytorch implementation of BRECQ, ICLR 2021

This is the codebase for the ICLR 2021 paper Trajectory Prediction using Equivariant Continuous Convolution