

It seems pretty obvious that an open weight model is a prime vector for social engineering/disinformation. If you can’t see the training data or mechanism, you are putting an awful lot of trust in the source. Better make sure that source has your best interests in mind.
















It seems like if you can recreate the training, the model could be considered open.
Lots of open source software performs operations on proprietary data, or accesses proprietary data/endpoints. That doesn’t make the software itself compromised.
LLMs are perhaps a little bit different in that the output (weights) is the part of the software that is useful.