← 返回论文检索
ICLR 2026PosterAccept (Poster)

A Probabilistic Hard Concept Bottleneck for Steerable Generative Models

María Martínez-García, Ricardo Vazquez Alvarez, Alejandro Lancho, Pablo Olmos, Isabel Valera

Saarland University, Universität des Saarlandes · Universidad Carlos III de Madrid · Saarland University, Saarland University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Concept Bottleneck Generative Models (CBGMs) incorporate a human-interpretable concept bottleneck layer, which makes them interpretable and steerable. However, designing such a layer for generative models poses the same challenges as for concept bottleneck models in a supervised context, if not greater ones. Deterministic mappings from the model inner representations to soft concepts in existing CBGMs: (i) limit steerable generation to modifying concepts in existing inputs; and, more importantly, (ii) are susceptible to *concept leakage*, which hinders their steerability. To address these limitations, we first introduce the Variational Hard Concept Bottleneck (VHCB) layer. The VHCB maps probabilistic estimates of binary latent variables to hard concepts, which have been shown to mitigate leakage. Remarkably, its probabilistic formulation enables direct generation from a specified set of concepts. Second, we propose a systematic evaluation framework for assessing the steerability of CBGMs across various tasks (e.g., activating and deactivating concepts). Our framework which allows us to empirically demonstrate that the VHCB layer consistently improves steerability.