Macro-F1 degrades more severely than Accuracy under open-world evaluation, revealing that classification failures disproportionately concentrate on behaviorally overlapping minority service classes. High Accuracy values in FreeNet (unknown services absorbed at 97.95% confidence) do not indicate successful detection of unseen services. Conventional single-metric reporting thus conceals the most operationally relevant failure modes.
From 2026-saleem-open-world-darknet-traffic — Open-World Darknet Traffic Recognition Under Leave-One-Service-Out Evaluation
· §IV-B, Fig. 3
· 2026
· arXiv preprint
Implications
Circumvention tool red-teaming should target Macro-F1 rather than Accuracy when benchmarking against traffic classifiers — high Accuracy can mask the complete failure to detect novel transport protocols if they are absorbed into existing service categories.
New protocol designs should be evaluated against leave-one-service-out classifiers, not closed-world baselines, to realistically estimate evasion durability.