Unsupervised Multi-Agent and Single-Agent Perception from Cooperative Views
Project page for UMS, accepted at CVPR 2026.
Abstract
UMS is an unsupervised framework that uses cooperative observations to learn both multi-agent and single-agent 3D object detection without human-annotated 3D boxes. Denser point clouds improve proposal classification, while geometric and semantic agreement across views supplies supervision to the single-agent branch.
Summary Video
Motivation
Method
UMS combines a Proposal Purifying Filter for removing unreliable candidate boxes, Progressive Proposal Stabilizing for easy-to-hard refinement, and Cross-View Consensus Learning for transferring multi-view geometric and semantic agreement to the single-agent detector.
Results
UMS achieves the strongest unsupervised results on both V2V4Real and OPV2V. At IoU 0.5, it reaches 52.03 AP for multi-agent and 44.27 AP for single-agent perception on V2V4Real; on OPV2V, it reaches 83.89 and 71.30 AP, respectively.
Ablation Study
The core ablation isolates Proposal Purifying Filter, Progressive Proposal Stabilizing, and Cross-View Consensus Learning, showing how the components jointly improve multi-agent and single-agent perception.
Qualitative Results
V2V4Real
BibTeX
@inproceedings{yang2026ums,
title = {Unsupervised Multi-Agent and Single-Agent Perception from Cooperative Views},
author = {Yang, Haochen and Li, Baolu and Li, Lei and Ren, Delin and Guo, Jiacheng and Qin, Minghai and Zhang, Tianyun and Yu, Hongkai},
booktitle = {IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}