Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Last update: Nov 25, 2022

Related tags

Deep Learning SGN

Overview

Semantic Grouping Network for Video Captioning

Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. AAAI 2021. [arxiv]

Environment

Ubuntu 16.04
CUDA 9.2
cuDNN 7.4.2
Java 8
Python 2.7.12
- PyTorch 1.1.0
- Other python packages specified in requirements.txt

Usage

1. Setup

$ pip install -r requirements.txt

2. Prepare Data

Download the GloVe Embedding from here and locate it at data/Embeddings/GloVe/GloVe_300.json.
Extract features from datasets and locate them at data/ /features/ .hdf5.

e.g. ResNet101 features of the MSVD dataset will be located at data/MSVD/features/ResNet101.hdf5.

I refer to this repo for extracting the ResNet101 features, and this repo for extracting the 3D-ResNext101 features.
Split the features into train, val, and test sets by running following commands.
```
$ python -m split.MSVD
$ python -m split.MSR-VTT
```

You can skip step 2-3 and download below files

MSVD
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]
MSR-VTT
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]

3. Prepare The Code for Evaluation

Clone the evaluation code from the official coco-evaluation repo.

$ git clone https://github.com/tylin/coco-caption.git
$ mv coco-caption/pycocoevalcap .
$ rm -rf coco-caption

4. Extract Negative Videos

$ python extract_negative_videos.py

or you can skip this step as the output files are already uploaded at data/ /metadata/neg_vids_ .json

5. Train

$ python train.py

You can change some hyperparameters by modifying config.py.

Pretrained Models - SGN(R101+RN)

*Disclaimer: The models above do not have the same weight as the models used in the paper (I trained them again because I lost).

6. Evaluate

$ python evaluate.py --ckpt_fpath

License

The source-code in this repository is released under MIT License.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Python 3 module to print out long strings of text with intervals of time inbetween

Time Series Cross-Validation -- an extension for scikit-learn

Official implementation for paper: Feature-Style Encoder for Style-Based GAN Inversion

A tutorial on DataFrames.jl prepared for JuliaCon2021

The code is an implementation of Feedback Convolutional Neural Network for Visual Localization and Segmentation.

Code release for "Making a Bird AI Expert Work for You and Me".

[3DV 2020] PeeledHuman: Robust Shape Representation for Textured 3D Human Body Reconstruction

Hypercomplex Neural Networks with PyTorch

PyTorch implementation of ECCV 2020 paper "Foley Music: Learning to Generate Music from Videos "

GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation. (CVPR 2021)

验证码识别深度学习 tensorflow 神经网络

Flower classification model that classifies flowers in 10 classes made using transfer learning (~85% accuracy).

A deep learning tabular classification architecture inspired by TabTransformer with integrated gated multilayer perceptron.

Explainable Medical ImageSegmentation via GenerativeAdversarial Networks andLayer-wise Relevance Propagation

Source code for our paper "Learning to Break Deep Perceptual Hashing: The Use Case NeuralHash"

Implementation of the SUMO (Slim U-Net trained on MODA) model

Face Recognize System on camera AI OAK1

Implementation of SegNet: A Deep Convolutional Encoder-Decoder Architecture for Semantic Pixel-Wise Labelling

AI Summer's complete catalog of articles

Official implementation of "A Unified Objective for Novel Class Discovery", ICCV2021 (Oral)

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Python 3 module to print out long strings of text with intervals of time inbetween

Time Series Cross-Validation -- an extension for scikit-learn

Official implementation for paper: Feature-Style Encoder for Style-Based GAN Inversion

A tutorial on DataFrames.jl prepared for JuliaCon2021

The code is an implementation of Feedback Convolutional Neural Network for Visual Localization and Segmentation.

Code release for "Making a Bird AI Expert Work for You and Me".

[3DV 2020] PeeledHuman: Robust Shape Representation for Textured 3D Human Body Reconstruction

Hypercomplex Neural Networks with PyTorch

PyTorch implementation of ECCV 2020 paper "Foley Music: Learning to Generate Music from Videos "

GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation. (CVPR 2021)

验证码识别 深度学习 tensorflow 神经网络

Flower classification model that classifies flowers in 10 classes made using transfer learning (~85% accuracy).

A deep learning tabular classification architecture inspired by TabTransformer with integrated gated multilayer perceptron.

Explainable Medical ImageSegmentation via GenerativeAdversarial Networks andLayer-wise Relevance Propagation

Source code for our paper "Learning to Break Deep Perceptual Hashing: The Use Case NeuralHash"

Implementation of the SUMO (Slim U-Net trained on MODA) model

Face Recognize System on camera AI OAK1

Implementation of SegNet: A Deep Convolutional Encoder-Decoder Architecture for Semantic Pixel-Wise Labelling

AI Summer's complete catalog of articles

Official implementation of "A Unified Objective for Novel Class Discovery", ICCV2021 (Oral)

验证码识别深度学习 tensorflow 神经网络