i see what your saying. i didnt mean to discredit standard benchmarks entirely.
i guess its obvious that it measures capability regardless of imprecision.
2 major proposed changes:
**first, i dont really know. aside from saying “benchmark your own prompt+usecase”
a proposed plan:
approach one: pay attention and credit new or improved architecture designs and research.
approach two: spend more attention on benchmarks. especially specific benchmarks ( that are not focused with industrial domain tasks.) **domain task pursuit, is useful!.. but it depends on if your interest align to popular domains.
approach three: if willing to utilize remotely hosted models. rating should also take in consideration… tools and everything else: websearch performance, RAG performance, smooth interface, pref/balance between speed vs comprehensiveness, cost (if relevant), etc… .
Honestly, I think the most reasonable approach is just to see what other people’s experience is like and which models are well regarded, then try them out and see which one is the best fit for what you’re doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.
i see what your saying. i didnt mean to discredit standard benchmarks entirely.
i guess its obvious that it measures capability regardless of imprecision.
2 major proposed changes:
**first, i dont really know. aside from saying “benchmark your own prompt+usecase”
a proposed plan:
Honestly, I think the most reasonable approach is just to see what other people’s experience is like and which models are well regarded, then try them out and see which one is the best fit for what you’re doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.