Comment on Automatically hiding spam works

  1. I suspect the archive post is not actually using the proper, technical definition of accuracy.

    Formally
    accuracy = (true positives + true negatives)/(total number of tests)
    In other words, if there are a millon works in the archive, 7000 would be misclassified, if the accuracy was actually 99.3%. As we do not know whether those 7000 represent false positives (legitimate works marked as spam) or false negatives (spam works not picked up), there is no way to figure out how many legitimate works would be hidden. We can speculate - as a few have done here :) - but the problem is that we simply do not know based on the information given.

    But as I said, I doubt we are ment to interpret it in the formal sense. I would be pretty interested in the false positive rate as well, if nothing else then from a purely stats fascination standpoint (as somebody who deals with the biostats of diagnostic test in my professional life ...)

    Last Edited Fri 24 Nov 2017 06:02PM UTC

    Comment Actions
    1. Stark Evolution in action

      I agree that it's unlikely AO3 means us to take the 99.3% in the most technical sense. I hope they mean that out of every 1000 works the Spam Detector *hides* (no knowing how many it checks) 7 of them could be incorrectly categorized as Spam.

      That would be a much lower number than .7% of all works and comments posted.

      Comment Actions