Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

On my LLM benchmarks Claude 3 Opus beats only one flavor of GPT-4: GPT-4 Turbo v3/1106-preview

All the other flavors are still better, with top winners: GPT-4 v1/0314 and GPT-4 Turbo v4/0125-preview.

The benchmark is based on prompts and tests from LLM-driven products, so it is biased towards business cases.



Most benchmarks use GPT4 as the grader. What does your benchmark use and do you believe that this causes any bias in the results?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: