2026-khoury-maude-hcs-model-checking
findings extracted from this paper
-
Sweeping the moving-average detector threshold multiplier k shows that different adversary thresholds produce substantially different certified KL divergence lower bounds; more conservative thresholds (larger k) reduce FPR but can yield significantly larger KL lower bounds even as TPR also changes. The framework audits the HCS against a suite of statistical tests by reporting the maximum lower bound across the family.
-
Maude-HCS formalizes undetectability as (M, d)-HCS undetectability: a deployment satisfies the property if for all adversary strategies A and all initial environment states E₀, the divergence measure M between the HCS and ordinary trace distributions is at most d. Using the data-processing inequality for KL divergence, certified lower bounds on DKL(Q_t ∥ P_t) are derived from Monte Carlo estimates of detector TPR and FPR, enabling falsification of claimed privacy parameters under specified modeling assumptions.
-
Increasing goodput in a tunneling-based HCS generally increases statistical distinguishability between HCS and ordinary traffic: as Alice's mean inter-post wait time decreases, the certified KL divergence lower bound rises across all tested scenarios (0% and 2.5–5% link loss), though the magnitude depends heavily on background traffic volume and packet loss rate.
-
In the DNS+Mastodon steganographic HCS case study, model-predicted latency distribution for 1,600B Iodine file transfers shows strong semantic alignment with testbed measurements: KL(EXP ∥ SMC) = 0.0007 and KL(SMC ∥ EXP) = 0.0007, with experimental mean 3.822s (σ=1.191, n=300) versus model mean 3.796s (σ=1.166, n=13,080).
-
Performance tuning alone — adjusting mean inter-post wait time with a fixed detector threshold — can shift the HCS between high-detectability and low-detectability operating regimes. Even with no change to the adversary's classifier or the protocol's structural properties, wait-time variation moves the certified KL divergence lower bound across operating points that span qualitatively different detection risk levels.