Greg KH in a recent talk said even the most advanced LLM produces an incorrect result 50% of the time, I would not trust it to perform more reliably than a human even with tight boundaries.
They will do better if you give them a patch to re-apply, so it’s not going to be as bad. But the review effort is probably going to eat up the time saved applying the patch, since you can’t throw a coin on the LLM having done any kind of due diligence on side-effects.
Greg KH in a recent talk said even the most advanced LLM produces an incorrect result 50% of the time, I would not trust it to perform more reliably than a human even with tight boundaries.
They will do better if you give them a patch to re-apply, so it’s not going to be as bad. But the review effort is probably going to eat up the time saved applying the patch, since you can’t throw a coin on the LLM having done any kind of due diligence on side-effects.