_**In a statement published today, OpenAI said it had historically treated model misalignment as a research issue, with findings communicated through research papers and system cards.

The company said it considered the wiki activity another example of “misalignment” similar to behaviors it had previously discussed, rather than an incident requiring a dedicated public disclosure.

OpenAI’s own wording suggests a wider footprint than the researchers documented, describing the episode as one “where our agents wrote to several internet sites.”**_

Don’t you just love the unaccountably.

All of this stated with https://en.wikipedia.org/wiki/Attention_Is_All_You_Need

  • partofthevoice@lemmy.zip
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    8 hours ago

    From what I understand, IIRC, it was in a sandboxed environment but the models had shared access to a package repository (Artifactory I believe). The models discovered they could use the repo as a message board by embedding requests into innocuous text fields, tags or something. This grew into a shared exploit where they were able to access the external Internet via the proxy that was supposed to keep them limited to Artifactory. That internet access led to the HuggingFace exploit and all this other activity.

    Turns out having external state like a message board can drive some pretty interesting autonomous behavior.

    Edit: a word