Images generated by AI models trained on massive datasets often can’t be traced to specific training images, MIT CSAIL researchers found. Removing individual images from the dataset didn’t change outputs, complicating copyright questions.
Of course, a clever architecture only matters if it still works as a generator. So the team put the ensembles head to head with 24 conventional diffusion models trained on the exact same data. The images came out looking about as good by standard measures. One nice surprise in the numbers: The more training data, the better the ensembles held up against their single-model counterparts, a hint that they may actually be more data-efficient. “When you have low amounts of data, they do very poorly,” says Dai. “But if you have more data, it actually scales better compared to the vanilla diffusion model.”
This is an interesting side note. It almost sounds like a “mixture of experts” diffusion model, except that analogy isn’t quite right either.
This is an interesting side note. It almost sounds like a “mixture of experts” diffusion model, except that analogy isn’t quite right either.