I suspect the archive post is not actually using the proper, technical definition of accuracy.
Formally
accuracy = (true positives + true negatives)/(total number of tests)
In other words, if there are a millon works in the archive, 7000 would be misclassified, if the accuracy was actually 99.3%. As we do not know whether those 7000 represent false positives (legitimate works marked as spam) or false negatives (spam works not picked up), there is no way to figure out how many legitimate works would be hidden. We can speculate - as a few have done here :) - but the problem is that we simply do not know based on the information given.
But as I said, I doubt we are ment to interpret it in the formal sense. I would be pretty interested in the false positive rate as well, if nothing else then from a purely stats fascination standpoint (as somebody who deals with the biostats of diagnostic test in my professional life ...)
I agree that it's unlikely AO3 means us to take the 99.3% in the most technical sense. I hope they mean that out of every 1000 works the Spam Detector *hides* (no knowing how many it checks) 7 of them could be incorrectly categorized as Spam.
That would be a much lower number than .7% of all works and comments posted.
Comment on Automatically hiding spam works
alrun Fri 24 Nov 2017 06:00PM UTC
Last Edited Fri 24 Nov 2017 06:02PM UTC
Comment Actions
AnonEhouse Sat 25 Nov 2017 12:49AM UTC
Comment Actions