Live inference with object interaction
≈18 ms online inference · ≤30 FPS camera limit
RTX 4080 Super · without ID tracking.
RGB-D · CATEGORY-LEVEL OBJECT POSE ESTIMATION
PIT IN Co., Republic of Korea
One full-frame predictor. Multiple objects.
9-DoF poses in real time.
RYOPO · RTX A6000
Mask-free variant · RTX A6000
No external instance segmentor
Accurate category-level pose estimation often starts with a separate instance segmentor, then processes each object crop. RYOPO brings detection, segmentation, and pose estimation into a single trainable RGB-D set predictor, without explicit CAD-derived shape priors.
Shared image and scene encoding avoids repeated crop encoding. Object queries connect measured 3D observations to an explicit pose state, refining rotation, translation, and metric size for multiple objects.
FROM CAPTURED DATA TO LIVE PREDICTIONS
Both demos use a RealSense L515 RGB-D camera and mask-free RYOPO models trained separately on their respective datasets.
The plush-toy model was trained on 12 videos of one physical plush toy. The two additional instances in the live recording were not included in its training data.
HEX-BOLT DEMONSTRATIONS
Hex-bolt annotations assume continuous symmetry about the local Y axis.
≈18 ms online inference · ≤30 FPS camera limit
RTX 4080 Super · without ID tracking.
Live RGB-D inference during rapid camera motion.
Left: per-frame RYOPO. Right: a separate tracker adds object IDs. This extension is not part of the paper’s benchmark results.
HOW WE LABEL THE CAPTURES
Our custom datasets were labeled with SnapPose, a browser-based tool for annotating object rotation, translation, and metric size in RGB-D point clouds.
Markers are used for annotation, not pose inference.
SnapPose on GitHub Public release planned · currently private
SEE THE PREDICTIONS
RYOPO, AG-Pose (DINO), CleanPose, and VI-Net on the same RGB-D frames.
20 FPS playback, not inference speed.
QUANTITATIVE RESULTS
Independent 3D IoU and all-object pose AP (%); missed detections count as failures.
| Method | IoU50 | 5° 2 cm | 5° 5 cm | 10° 5 cm |
|---|---|---|---|---|
| DPDN | 83.4 | 38.1 | 42.2 | 65.3 |
| HS-Pose | 82.8 | 38.5 | 46.6 | 70.1 |
| VI-Net | 83.1 | 40.7 | 47.1 | 68.3 |
| AG-Pose (DINO) | 84.1 | 46.4 | 52.7 | 71.0 |
| CleanPose | 84.1 | 50.4 | 55.1 | 72.3 |
| RYOPO | 93.0 | 38.7 | 51.1 | 73.6 |
| RYOPO, mask-free | 91.9 | 36.7 | 48.5 | 71.9 |
Per-object methods use the benchmark’s supplied predicted masks; RYOPO uses its own predictions.
Mask-free: separately trained with adjusted settings. Accuracy: eager FP32.
HOW IT WORKS
Built on the DETR-style EdgeCrafter backbone, RYOPO integrates object-associated depth observations and explicit pose-state refinement into a shared full-frame predictor.
Predicted object regions associate measured depth points with each query. Geometry, image features, and observation support form object point descriptors.
Object and shared scene features update the queries during set decoding, connecting image-based predictions to 3D observations.
The object query and current pose condition attention to point descriptors. Residual corrections update the pose and condition subsequent refinement.
Provisional entry. The arXiv ID will be added after submission.
@misc{lee2026ryopo,
title = {{RYOPO}: Bringing End-to-End Category-Level Object Pose Estimation into Real Time},
author = {Hakjin Lee and Junghoon Seo and Jaehoon Sim},
year = {2026},
eprint = {ARXIV_ID_PENDING},
archivePrefix = {arXiv}
}