A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generation
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Controllable hand image generation aims to synthesize geometrically accurate images with consistent appearance. Recently, diffusion models have been widely applied for hand image synthesis. However, through input-level fusion or feature-level modulation, existing methods inject control signals with fixed strength across all timesteps, ignoring the progressive nature of the denoising process. In this paper, we reveal that the modulation of control signals depends on the denoising state and condition complexity. Due to distinct semantic distributions and information densities, achieving effective interaction among these heterogeneous representations remains challenging. To address this, we propose a Temporal and Content Co-Awareness Latent Diffusion method that introduces a dual-driven modulation strategy. Specifically, we design a query-based interaction mechanism to mitigate information redundancy and align semantic distributions. Leveraging cross-domain interaction, the model infers required control information to dynamically adjust pose and appearance injection strengths. Furthermore, we design a Pose-Invariant Appearance Encoder that captures both global appearance consistency and local texture details. Extensive experiments validate our superiority over state-of-the-art.