I3D

简介

@inproceedings{inproceedings,
  author = {Carreira, J. and Zisserman, Andrew},
  year = {2017},
  month = {07},
  pages = {4724-4733},
  title = {Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset},
  doi = {10.1109/CVPR.2017.502}
}

@article{NonLocal2018,
  author =   {Xiaolong Wang and Ross Girshick and Abhinav Gupta and Kaiming He},
  title =    {Non-local Neural Networks},
  journal =  {CVPR},
  year =     {2018}
}

模型库

Kinetics-400

配置文件	分辨率	GPU 数量	主干网络	预训练	top1 准确率	top5 准确率	推理时间 (video/s)	GPU 显存占用 (M)	ckpt	log	json
i3d_r50_32x2x1_100e_kinetics400_rgb	340x256	8	ResNet50	ImageNet	72.68	90.78	1.7 (320x3 frames)	5170	ckpt	log	json
i3d_r50_32x2x1_100e_kinetics400_rgb	短边 256	8	ResNet50	ImageNet	73.27	90.92	x	5170	ckpt	log	json
i3d_r50_video_32x2x1_100e_kinetics400_rgb	短边 256p	8	ResNet50	ImageNet	72.85	90.75	x	5170	ckpt	log	json
i3d_r50_dense_32x2x1_100e_kinetics400_rgb	340x256	8x2	ResNet50	ImageNet	72.77	90.57	1.7 (320x3 frames)	5170	ckpt	log	json
i3d_r50_dense_32x2x1_100e_kinetics400_rgb	短边 256	8	ResNet50	ImageNet	73.48	91.00	x	5170	ckpt	log	json
i3d_r50_lazy_32x2x1_100e_kinetics400_rgb	340x256	8	ResNet50	ImageNet	72.32	90.72	1.8 (320x3 frames)	5170	ckpt	log	json
i3d_r50_lazy_32x2x1_100e_kinetics400_rgb	短边 256	8	ResNet50	ImageNet	73.24	90.99	x	5170	ckpt	log	json
i3d_nl_embedded_gaussian_r50_32x2x1_100e_kinetics400_rgb	短边 256p	8x4	ResNet50	ImageNet	74.71	91.81	x	6438	ckpt	log	json
i3d_nl_gaussian_r50_32x2x1_100e_kinetics400_rgb	短边 256p	8x4	ResNet50	ImageNet	73.37	91.26	x	4944	ckpt	log	json
i3d_nl_dot_product_r50_32x2x1_100e_kinetics400_rgb	短边 256p	8x4	ResNet50	ImageNet	73.92	91.59	x	4832	ckpt	log	json

注：

这里的 GPU 数量 指的是得到模型权重文件对应的 GPU 个数。默认地，MMAction2 所提供的配置文件对应使用 8 块 GPU 进行训练的情况。依据线性缩放规则，当用户使用不同数量的 GPU 或者每块 GPU 处理不同视频个数时，需要根据批大小等比例地调节学习率。如，lr=0.01 对应 4 GPUs x 2 video/gpu，以及 lr=0.08 对应 16 GPUs x 4 video/gpu。
这里的 推理时间 是根据基准测试脚本获得的，采用测试时的采帧策略，且只考虑模型的推理时间，并不包括 IO 时间以及预处理时间。对于每个配置，MMAction2 使用 1 块 GPU 并设置批大小（每块 GPU 处理的视频个数）为 1 来计算推理时间。
我们使用的 Kinetics400 验证集包含 19796 个视频，用户可以从验证集视频下载这些视频。同时也提供了对应的数据列表（每行格式为：视频 ID，视频帧数目，类别序号）以及标签映射（类别序号到类别名称）。

对于数据集准备的细节，用户可参考数据集准备文档中的 Kinetics400 部分。

如何训练

用户可以使用以下指令进行模型训练。

python tools/train.py ${CONFIG_FILE} [optional arguments]

例如：以一个确定性的训练方式，辅以定期的验证过程进行 I3D 模型在 Kinetics400 数据集上的训练。

python tools/train.py configs/recognition/i3d/i3d_r50_32x2x1_100e_kinetics400_rgb.py \
    --work-dir work_dirs/i3d_r50_32x2x1_100e_kinetics400_rgb \
    --validate --seed 0 --deterministic

更多训练细节，可参考基础教程中的 训练配置 部分。

如何测试

用户可以使用以下指令进行模型测试。

python tools/test.py ${CONFIG_FILE} ${CHECKPOINT_FILE} [optional arguments]

例如：在 Kinetics400 数据集上测试 I3D 模型，并将结果导出为一个 json 文件。

python tools/test.py configs/recognition/i3d/i3d_r50_32x2x1_100e_kinetics400_rgb.py \
    checkpoints/SOME_CHECKPOINT.pth --eval top_k_accuracy mean_class_accuracy \
    --out result.json --average-clips prob

更多测试细节，可参考基础教程中的 测试某个数据集 部分。

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

README_zh-CN.md

README_zh-CN.md

I3D

简介

模型库

Kinetics-400

如何训练

如何测试

Files

README_zh-CN.md

Latest commit

History

README_zh-CN.md

File metadata and controls

I3D

简介

模型库

Kinetics-400

如何训练

如何测试