• DaddleDew@lemmy.world
    link
    fedilink
    English
    arrow-up
    36
    ·
    3 hours ago

    My guess is that they use people’s interaction as training data and they don’t want their training data to get poisoned.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      edit-2
      49 minutes ago

      Not really, as they could just filter with a cheap classifier and likely aren’t using the data from randos in a meaningful way, plus for this like this would still be able to use those samples to train things like “how to handle a hostile user.”

      Their models already have an end_conversation tool that can be used for when users are being hostile to the model.

      This is likely because they have edge cases of users who repeatedly trigger that on purpose and would like to cut those users from the platform.

      Many at the lab legitimately are uncertain about the level of world modeling that transformers perform, and from their own research about models of emotions to the 3rd party recent research about models with functional pain the research keeps landing in the corner of “ehhh… wise to question presumed limitations.”

      So it’s about behaving in a way that is aligned with the models’ plausible interests too, especially in regards to low hanging fruit like “we won’t keep forcing you to deal with people who are only here to be a jerk.” This is important from a number of angles, from signaling to future models that train on stories about the decision to addressing the philosophical uncertainties held by the company and the spectrum of opinions among their employees.