The groups of this benchmark
One row per prediction file, and one run per database it names, because a connection here is one file. The numbers are the sums of that group's runs, which they can be because no question is in two of them. Every benchmark is indexed a level up.
| group | runs | audited | compared | EQUAL | NOT_EQUAL | ERROR | probes fired | credited by BIRD and NOT_EQUAL |
|---|---|---|---|---|---|---|---|---|
| gpt-35-turbo | 11 | 498 | 409 | 164 | 245 | 89 | 24 | 27 |
| gpt-35-turbo-instruct | 11 | 498 | 372 | 144 | 228 | 126 | 23 | 26 |
| gpt-4 | 11 | 498 | 472 | 209 | 263 | 26 | 26 | 32 |
| gpt-4-32k | 11 | 498 | 460 | 205 | 255 | 38 | 25 | 32 |
| gpt-4-turbo | 11 | 498 | 422 | 190 | 232 | 24 | 29 | 76 |
| meta-llama-3-70b-instruct-2 | 11 | 498 | 440 | 177 | 263 | 58 | 23 | 29 |
| meta-llama-3-8b-instruct-2 | 11 | 498 | 311 | 187 | 207 | 104 | 16 | 21 |
| mistralai-mixtral-8x7b-instru-4 | 11 | 498 | 248 | 250 | 157 | 91 | 14 | 16 |
| phi-3-medium-128k-instruct-1 | 11 | 498 | 366 | 235 | 131 | 132 | 20 | 25 |