Paper overview
Abstract
Cloud-edge collaborative decoding keeps private data on the edge while combining predictions from an edge small language model with a cloud large language model. However, the signals exchanged during decoding can still reveal private-context information. This work introduces an evaluation framework based on constructed question-answering datasets to audit this leakage, and proposes CoVeil, a decoding-time defense that dynamically optimizes transmitted signals to suppress leakage while preserving collaborative utility. Across the reported evaluations, CoVeil reduces data leakage by up to 87.2% with minimal accuracy loss.
