FlowMoDE: Coarse-to-fine Flow Matching for Structure-aware Sim-to-Real Monocular Depth Estimation

Liangjing Shao1,2, Wanhao Liu2, Jinsong Lin1, Zhiwei Fang1, Hongliang Ren1,2,#,

1 Department of Electronic Engineering, The Chinese University of Hong Kong
2 EACV Research Center, Shenzhen Loop Area Institute
Submitted to ICRA 2027

# Corresponding Author

Monocular RGB to Depth Map to Point Cloud: Baxter & Franka & Kuka

Abstract

Monocular depth estimation is a fundamental capability for 3D spatial perception in robotic and intelligent systems. While data‑driven discriminative methods have achieved compelling performance for monocular depth estimation, generative approaches can reconstruct fine structural details with stronger generalization. However, most existing generative diffusion models suffer from low inference efficiency caused by multi‑step sampling, whereas single‑step diffusion models struggle to recover high‑fidelity depth maps. To mitigate these drawbacks, this paper presents FlowMoDE, a coarse‑to‑fine flow‑matching‑based framework for efficient, structure‑aware monocular depth estimation supporting sim‑to‑real and cross‑scene generalization. The core architecture first conducts coarse‑level decoding conditioned on diffused features extracted from a pre‑trained semantic encoder, and subsequently restores fine‑grained structural patterns via expanded decoding layers. Trained exclusively on a single synthetic indoor dataset, our proposed method outperforms prior methods across five real‑world indoor and outdoor datasets, yielding an average 14.1% reduction in relative error against state‑of‑the‑art works as well as 9.9% reduction in chamfer distance for point cloud reconstruction. We further validate the generalization capability of FlowMoDE on realistic robotic scenes, with qualitative comparisons against SOTA methods.

Qualitative Comparison on Robotic Scene with Baxter

Reference result
FlowMoDE result

Qualitative Comparison on Robotic Scene with Franka

Reference result
FlowMoDE result

BibTeX

@article{shao2026flowmode,
  title={FlowMoDE: Coarse-to-fine Flow Matching for Structure-aware Sim-to-Real Monocular Depth Estimation},
  author={Liangjing Shao and Wanhao Liu and Jinsong Lin and Zhiwei Fang and Hongliang Ren},
  journal={Submitting},
  year={2026},
  url={coming soon}
}