Comment on Automatically hiding spam works

  1. I think your premise is flawed. According to your math, of 1000 legitimate works, 7 will be accidentally hidden while 993 remain visible. It isn't 7 legitimate works hidden for every 1 spam work hidden, it's 993 spam works hidden for every 7 legitimate works hidden.

    That percentage is just a best guess; it could be better or worse than that in reality.

    Comment Actions
    1. You failed to consider the "1000 legitimate works for every spam work" part, which is kind of a big deal since the problem I described is typical of situations with a relatively low percent of true positives.

      Comment Actions
      1. The number of legitimate/spam works is not relevant, the 0,7% false positive is of works already hidden and not works in general.

        So, of all the works in the archive, a given number of works will be hidden because the AutoMagic system will assume it is spam. Out of this number of works that the system flagged, 0,7% will be false positives. The 0,7% is of the works assumed to be spam by the system and not of all the works.

        So, using your number for an example, "1000 legitimate works for every spam work":

        I you assume 1 million works on AO3, about 1000 of those would be spam and flagged as so by the automatic system. Between those works flagged and hidden, there will be a 0,7% of false positives hidden incorrectly (meaning 7 legitimate works in this case).

        So what you would have is, roughly, 7 legitimate works hidden by mistake, 993 spam correctly flagged and hidden, and 999000 works the system ignored because it recognized it as not spam and didn't touch it.

        Comment Actions
        1. And a lot more time for Abuse to address the 7 works hidden by mistake (rather than addressing the 993 spam works).

          Comment Actions
        2. >the 0,7% false positive is of works already hidden and not works in general.

          That is not what false positive means.
          The meaning of false positive is probability of a positive on test A conditional to A being negative.
          What you are using is probability of A being negative conditional to a positive on test A.

          Nobody uses your definition because it is very sensitive to the ratio of positive to negative in the population you are testing, while the correct one is only sensitive to the quality of the testing process.

          If you want to make up your own flawed definitions, don't go around "correcting" people that follow the norm, darling.

          Comment Actions
          1. (Previous comment deleted.)

            1. >"I can't read", the comment

              Anon, you have to go back

              Comment Actions
            2. holy FUCK i aspire to this level of sarcasm this is seriously impressive. #goals

              "woah, we have a specialist!" "because I'm not a pedantic asshole" i love this

              Comment Actions
      2. Slightly higher than 99% is a low percent?

        I think what's confusing me is the number 0,7%. The only percentage we know is: "99.3% accuracy rate when it comes to identifying spam works and comments". That remaining 0,7% =/= false positives, because it has to include false negatives as well, and it counts comment as well as works, so if you have 1000 legitimate works it's likely that less than 7 will be hidden, irregardless of the huge number of works being correctly marked as spam.

        I don't understand how you get "the system will hide 7 legitimate works for every spam work hidden". It's that ratio of 7 legitimate works hidden per 1 spam work hidden, when legitimate works hidden is the far smaller percentage. Wouldn't it be as the other are pointing out, something more like for every 7 legitimate works hidden, 7000 spam works are correctly hidden? That is, isn't the inverse true? (Maths are not my specialty, and googling false positive in relation to spam gave me articles like: https://www.networkworld.com/article/2327896/lan-wan/what-is-a-false-positive-.html which are for me hard to parse.)

        It's not great that some works (really, likely crack and people putting links to suspect websites in their work) will be hidden, it will be a low number, easily rectified, as opposed to the problem now of having humans slowly go through everything?

        Comment Actions
        1. That number is not especially important, as it is the 1000 legit per 1 spam ratio, they are there only to show how even very favourable numbers can result in poor performance.

          I chose 0.7% because usually false positive ratio and false negative ratio are not too different. Maybe one is 5, 10 times the other, but rarely more.

          The rest of your comment shows a lack of knowledge in the field of statistics: it will be hard for you to understand the issue you're stuck on without conditional probability at the very least.

          Let's make it even simpler anyways.

          You have 10010 works. Of those works, 10 are spam, 10000 are legitimate.
          You have a very good spam system thay always catches spam works (0% false negative rate) and only mistakes 1 legit work for spam ofut of 100 legit works (1% false positive rate).
          Your system will find all 10 real spam works, and will confuse 1% of 10000 legit works as spam. 1% of 10000 is 100.
          In conclusion, your system will claim to have found 110 spam works, of which only 10 are spam: that means only 1 in 11 works marked as spam will be actual spam.

          The problem gets worse the more legit works there are for each spam work, and the higher the false positive rate.
          Clear enough?

          Comment Actions
    2. Green signpost saying Adopt-a-Highway: Butterflies & Rainbows

      While I agree that OP's math is wrong, it would be nice to know the ratio of legitamate works to spam works. If the ratio of spam works to legitamate ones posted every day is one-to-one, seven legitamate works woud be hidden per thousand legitamate works—but if the spammers are posting more spam works than legitimate works get posted every day, then the number of incorrectly hidden works per thousand would go up, and the inverse is true.

      It'd be nice to guess around how many legitimate works per thousand are getting hidden, which you can't do with the current information without also knowing an average ratio of spam works to legitamate works posted.

      Comment Actions
      1. You should be more wary of anonymous commenters that misunderstand basic statistical terms.

        For reference: https://en.wikipedia.org/wiki/False_positive_rate

        Comment Actions
        1. Green signpost saying Adopt-a-Highway: Butterflies & Rainbows

          Oh shoot, you're right. I read their comment and thought what the archive had said was the ratio of spam works to legitamate works hidden instead of the plain accuracy rate. I should have reread the update before posting. Thanks for the correction.

          Comment Actions
    3. I suspect the archive post is not actually using the proper, technical definition of accuracy.

      Formally
      accuracy = (true positives + true negatives)/(total number of tests)
      In other words, if there are a millon works in the archive, 7000 would be misclassified, if the accuracy was actually 99.3%. As we do not know whether those 7000 represent false positives (legitimate works marked as spam) or false negatives (spam works not picked up), there is no way to figure out how many legitimate works would be hidden. We can speculate - as a few have done here :) - but the problem is that we simply do not know based on the information given.

      But as I said, I doubt we are ment to interpret it in the formal sense. I would be pretty interested in the false positive rate as well, if nothing else then from a purely stats fascination standpoint (as somebody who deals with the biostats of diagnostic test in my professional life ...)

      Last Edited Fri 24 Nov 2017 06:02PM UTC

      Comment Actions
      1. Stark Evolution in action

        I agree that it's unlikely AO3 means us to take the 99.3% in the most technical sense. I hope they mean that out of every 1000 works the Spam Detector *hides* (no knowing how many it checks) 7 of them could be incorrectly categorized as Spam.

        That would be a much lower number than .7% of all works and comments posted.

        Comment Actions