← 返回论文检索
NeurIPS 2025{location} PosterAccept (poster)

PUO-Bench: A Panel Understanding and Operation Benchmark with A Privacy-Preserving Framework

Wei LIN, Yiwei Zhou, Junkai Zhang, Rui Shao, Zhiyuan Zhao, Junyu Gao, Antoni Chan, Xuelong Li

City University of Hong Kong · Beijing Institute of Technology · Tsinghua University · Northwest Polytechnical University Xi'an · China Telecom · Northwestern Polytechnical University, Center for OPTical IMagery Analysis and Learning

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Recent advancements in Vision-Language Models (VLMs) have enabled GUI agents to leverage visual features for interface understanding and operation in the digital world. However, limited research has addressed the interpretation and interaction with control panels in real-world settings. To bridge this gap, we propose the Panel Understanding and Operation (PUO) benchmark, comprising annotated panel images from appliances and associated vision-language instruction pairs. Experimental results on the benchmark demonstrate significant performance disparities between zero-shot and fine-tuned VLMs, revealing the lack of PUO-specific capabilities in existing language models. Furthermore, we introduce a Privacy-Preserving Framework (PPF) to address privacy concerns in cloud-based panel parsing and reasoning. PPF employs a dual-stage architecture, performing panel understanding on edge devices while delegating complex reasoning to cloud-based LLMs. Although this design introduces a performance trade-off due to edge model limitations, it eliminates the transmission of raw visual data, thereby mitigating privacy risks. Overall, this work provides foundational resources and methodologies for advancing interactive human-machine systems and robotic field in panel-centric applications.