TY - GEN
T1 - VoxHammer
T2 - 13th International Conference on 3D Vision, 3DV 2026
AU - Li, Lin
AU - Huang, Zehuan
AU - Feng, Haoran
AU - Zhuang, Gengxiong
AU - Chen, Rui
AU - Guo, Chunchao
AU - Sheng, Lu
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - 3D local editing of specified regions is crucial for the game industry and robot interaction. Recent methods typically edit rendered multi-view images and then reconstruct 3D models, but they face challenges in precisely preserving unedited regions and overall coherence. Inspired by structured 3D generative models, we propose VoxHammer, a novel training-free approach that performs precise and coherent editing in 3D latent space. Given a 3D model, VoxHammer first predicts its inversion trajectory and obtains its inverted latents and key-value tokens at each timestep. Subsequently, in the denoising and editing phase, we replace the denoising features of preserved regions with the corresponding inverted latents and cached key-value tokens. By retaining these contextual features, this approach ensures consistent reconstruction of preserved areas and coherent integration of edited parts. To evaluate the consistency of preserved regions, we constructed Edit3D-Bench, a human-annotated dataset comprising hundreds of samples, each with carefully labeled 3D editing regions. Experiments demonstrate that VoxHammer significantly outperforms existing methods in terms of both 3D consistency of preserved regions and overall quality. Our method holds promise for synthesizing high-quality edited paired data, thereby laying the data foundation for in-context 3D generation.
AB - 3D local editing of specified regions is crucial for the game industry and robot interaction. Recent methods typically edit rendered multi-view images and then reconstruct 3D models, but they face challenges in precisely preserving unedited regions and overall coherence. Inspired by structured 3D generative models, we propose VoxHammer, a novel training-free approach that performs precise and coherent editing in 3D latent space. Given a 3D model, VoxHammer first predicts its inversion trajectory and obtains its inverted latents and key-value tokens at each timestep. Subsequently, in the denoising and editing phase, we replace the denoising features of preserved regions with the corresponding inverted latents and cached key-value tokens. By retaining these contextual features, this approach ensures consistent reconstruction of preserved areas and coherent integration of edited parts. To evaluate the consistency of preserved regions, we constructed Edit3D-Bench, a human-annotated dataset comprising hundreds of samples, each with carefully labeled 3D editing regions. Experiments demonstrate that VoxHammer significantly outperforms existing methods in terms of both 3D consistency of preserved regions and overall quality. Our method holds promise for synthesizing high-quality edited paired data, thereby laying the data foundation for in-context 3D generation.
KW - 3d diffusion model
KW - 3d editing
KW - 3d generation
UR - https://www.scopus.com/pages/publications/105042059543
U2 - 10.1109/3DV69130.2026.00126
DO - 10.1109/3DV69130.2026.00126
M3 - 会议稿件
AN - SCOPUS:105042059543
T3 - Proceedings - 2026 International Conference on 3D Vision, 3DV 2026
SP - 1281
EP - 1292
BT - Proceedings - 2026 International Conference on 3D Vision, 3DV 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 20 March 2026 through 23 March 2026
ER -