FINDING · EVALUATION

Macro-F1 degrades more severely than Accuracy under open-world evaluation, revealing that classification failures disproportionately concentrate on behaviorally overlapping minority service classes. High Accuracy values in FreeNet (unknown services absorbed at 97.95% confidence) do not indicate successful detection of unseen services. Conventional single-metric reporting thus conceals the most operationally relevant failure modes.

From 2026-saleem-open-world-darknet-trafficOpen-World Darknet Traffic Recognition Under Leave-One-Service-Out Evaluation · §IV-B, Fig. 3 · 2026 · arXiv preprint

Implications

Tags

censors
generic
techniques
ml-classifiertraffic-shape

Extracted by claude-sonnet-4-6 — review before relying.