← 返回论文检索
ACM Multimedia 2025Experience: Multimedia Applications

Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition Estimation

Donglin Zhang 0001, Boyuan Ma, Xiaojun Wu 0001, Josef Kittler

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754588 ↗

摘要

Food plays a vital role in human health, and accurate nutrition estimation is crucial for guiding healthy dietary choices. Traditional biochemical-based assessment methods are often inefficient, costly, and impractical for daily use. With the continuous progress in computer vision, some vision-based nutrition estimation approaches have emerged, typically relying on RGB images alone or in combination with depth images to infer nutritional information. These methods have achieved promising performance and garnered considerable attention. However, these methods often ignore visually imperceptible ingredients such as oil, sugar, and salt, which may significantly influence the estimation of nutritional content. Besides, existing methods lack explicit mechanisms for modeling nutrient-specific information and guiding attention toward nutrition-relevant semantics. To solve the above two issues, we propose a novel ingredients-guided and nutrients-prompted nutrition estimation method. Our method adopts multi-scale feature fusion and integrates RGB and depth modalities to enhance visual representation learning. To account for invisible ingredients, we introduce an ingredients-guided strategy, which enhances the sensitivity to non-visible nutritional factors. Moreover, a nutrient-prompt mechanism is introduced to explicitly guide the focus of the model toward nutrient-relevant attributes during estimation. We validate our method on Nutrition5k, where it consistently outperforms existing state-of-the-art methods, demonstrating its efficacy.