5 ms·
This is just typical of so much work in the field. They pick and choose which models to compare against and on which benchmarks. If this model was truly great,
by simonhughes22 3y ago
This is just typical of so much work in the field. They pick and choose which models to compare against and on which benchmarks. If this model was truly great, they would be comparing against Claude 2 and GPT4 across a bunch of different benchmarks. Instead they compare against Palm 2, which in a lot of tests is a weak model (https://venturebeat.com/ai/google-bard-fails-to-deliver-on-its-promise-even-after-latest-updates/#:~:text=The%20crux%20of%20the%20problem,content%20it%20has%20been%20fed https://venturebeat.com/ai/google-bard-fails-to-deliver-on-i....) and prone to hallucination (https://github.com/vectara/hallucination-leaderboard https://github.com/vectara/hallucination-leaderboard).