RGB-D · CATEGORY-LEVEL OBJECT POSE ESTIMATION

RYOPOBringing End-to-End Category-Level
Object Pose Estimation into Real Time

Hakjin LeeJunghoon SeoJaehoon Sim

PIT IN Co., Republic of Korea

One full-frame predictor. Multiple objects.
9-DoF poses in real time.

Category-level predictionREAL275
Mask-free live demonstrationRGB-D camera
31.8 FPS

RYOPO · RTX A6000

51.4 FPS

Mask-free variant · RTX A6000

RGB + Depth

No external instance segmentor

From a scene to
object poses.

Accurate category-level pose estimation often starts with a separate instance segmentor, then processes each object crop. RYOPO brings detection, segmentation, and pose estimation into a single trainable RGB-D set predictor, without explicit CAD-derived shape priors.

Shared image and scene encoding avoids repeated crop encoding. Object queries connect measured 3D observations to an explicit pose state, refining rotation, translation, and metric size for multiple objects.

OVERVIEW

FROM CAPTURED DATA TO LIVE PREDICTIONS

Real-world showcases

Both demos use a RealSense L515 RGB-D camera and mask-free RYOPO models trained separately on their respective datasets.

The plush-toy model was trained on 12 videos of one physical plush toy. The two additional instances in the live recording were not included in its training data.

Training captures12 videos · One instance
Cropped training views · 5 FPS playback.
Live predictionsSeen + two unseen
Original recording · 30 FPS playback.

HOW WE LABEL THE CAPTURES

RGB-D annotation with SnapPose

Our custom datasets were labeled with SnapPose, a browser-based tool for annotating object rotation, translation, and metric size in RGB-D point clouds.

  1. Capture RGB-D with marker-based camera poses.
  2. Label a reference frame in the 3D workspace.
  3. Propagate and inspect labels across a static-object capture using camera poses.

Markers are used for annotation, not pose inference.

SnapPose on GitHub Public release planned · currently private

Default example: NOCS REAL275, manually annotated with SnapPose.

SEE THE PREDICTIONS

Benchmark comparisons

RYOPO, AG-Pose (DINO), CleanPose, and VI-Net on the same RGB-D frames.

Predicted cuboidsGround truthRGB axes: x / y / z

20 FPS playback, not inference speed.

QUANTITATIVE RESULTS

All-object accuracy

Independent 3D IoU and all-object pose AP (%); missed detections count as failures.

REAL275 all-object accuracy. Higher is better.
MethodIoU505° 2 cm5° 5 cm10° 5 cm
DPDN83.438.142.265.3
HS-Pose82.838.546.670.1
VI-Net83.140.747.168.3
AG-Pose (DINO)84.146.452.771.0
CleanPose84.150.455.172.3
RYOPO93.038.751.173.6
RYOPO, mask-free91.936.748.571.9

Per-object methods use the benchmark’s supplied predicted masks; RYOPO uses its own predictions.

Mask-free: separately trained with adjusted settings. Accuracy: eager FP32.

HOW IT WORKS

Object queries meet measured geometry.

Built on the DETR-style EdgeCrafter backbone, RYOPO integrates object-associated depth observations and explicit pose-state refinement into a shared full-frame predictor.

(A)

Query-conditioned geometry

Predicted object regions associate measured depth points with each query. Geometry, image features, and observation support form object point descriptors.

(B)

Geometry feedback

Object and shared scene features update the queries during set decoding, connecting image-based predictions to 3D observations.

(C)

Pose-state refinement

The object query and current pose condition attention to point descriptors. Residual corrections update the pose and condition subsequent refinement.

Citation

Provisional entry. The arXiv ID will be added after submission.

@misc{lee2026ryopo,
  title         = {{RYOPO}: Bringing End-to-End Category-Level Object Pose Estimation into Real Time},
  author        = {Hakjin Lee and Junghoon Seo and Jaehoon Sim},
  year          = {2026},
  eprint        = {ARXIV_ID_PENDING},
  archivePrefix = {arXiv}
}

Architecture

Scroll horizontally on small screens to inspect the diagram. Press Escape to close.