q1130

R-SET NOT_EQUAL

european_football_2 · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json

The question

What are the short name of team who played safe while creating chance of passing?

the hint the set supplies: played safe while creating chance of passing refers to chanceCreationPassingClass = 'Safe'; short name of team refers to team_short_name

NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.

read by hand, 2026-09-04: B harmless: one distinct row repeated; DISTINCT added by the prediction; a list under twice the answer with every row present

43 distinct names, 44 or 56 rows

a maintainer's reading of this question, out of classification.json, copied from plans/reports/prediction-mode-260904-real-predictions/classification.json. It is not a verdict and nothing above it was computed from it.

The statements

gold

SELECT DISTINCT t1.team_short_name FROM Team AS t1 INNER JOIN Team_Attributes AS t2 ON t1.team_api_id = t2.team_api_id WHERE t2.chanceCreationPassingClass = 'Safe'

this statement states no ordering of its own

sha256:c276a97fdd20778fc5abe17922ef7d3f2e6eb13f75247763335e148959bb3166

second

SELECT Team.team_short_name
FROM Team
JOIN Team_Attributes ON Team.team_api_id = Team_Attributes.team_api_id
WHERE Team_Attributes.chanceCreationPassingClass = 'Safe'

this statement states no ordering of its own

sha256:00e1ae76585ce49c99428a6eae93166d84cbf7ad4afa7ecbc7c6747eea11e5ee

The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.

The rows they differ in

in gold, not in the prediction, 0 rows

from counterexample.json, 0 rows, up to 25 shown per side

side team_short_nametext
no rows

in the prediction, not in gold, 13 rows

from counterexample.json, 13 rows, up to 25 shown per side, 9 on this page

sidetimes team_short_nametext
second2 ARS
second2 BOL
second2 GEN
second2 WIS
second1 CAG
second1 LIV
second1 NAC
second1 REG
second1 WHU

What the benchmark would have said

BIRD's own check: 1

set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal

https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py

gold_rows
43
second_rows
56
gold_distinct_rows
43
second_distinct_rows
43

the test-suite check: 0

result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them

ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py

order_matters
false
gold_rows
43
second_rows
56
gold_columns
1
second_columns
1

the class: multiplicity

The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.

how many rows each result holds, and how many rows both hold, from evidence-gold.json and evidence-second.json43 gold56 prediction43 rows both results hold
The gold returned 43 rows and the prediction 56. 43 rows occur in both results at least once.
gold_types
["text"]
second_types
["text"]
multiset_equal
false
set_equal
true
order_equal
false
shorter_result_is_a_prefix
false

The results

gold, 43 rows

from counterexample.json, 43 rows, 25 on this page

team_short_nametext
WHU
HUE
BET
ROD
NAC
BAR
SAM
WAA
EMP
HAA
UTR
UDI
NAP
ARL
BMU
WIS
ARK
COR
REG
STK
SPA
PSV
FRE
CAG
ZAG
second, 56 rows

from counterexample.json, 56 rows, 25 on this page

team_short_nametext
HAA
ARK
ARL
ARS
ARS
ARS
BAR
BMU
BOL
BOL
BOL
BRE
CAG
CAG
CAT
COR
DUF
EMP
COT
UTR
FRE
FRO
GEN
GEN
GEN

The probes

A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.

ordering-over-numeric-text not applicable

this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10

the statement states no top level ORDER BY

what it measured
{
  "heuristic": true,
  "reason": "the statement states no top level ORDER BY"
}

arbitrary-cut not applicable

this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero

the statement states no LIMIT

what it measured
{
  "heuristic": true,
  "reason": "the statement states no LIMIT"
}

not-a-function-of-the-data quiet

rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data

what it measured
{
  "heuristic": true,
  "rule": "R-SET",
  "baseline_result_hash": "sha256:c276a97fdd20778fc5abe17922ef7d3f2e6eb13f75247763335e148959bb3166",
  "baseline_result": {
    "columns": [
      {
        "name": "team_short_name",
        "declared_type": "text"
      }
    ],
    "row_count": 43,
    "truncated": false,
    "rows_shown": 10,
    "rows": [
      [
        {
          "type": "str",
          "value": "WHU"
        }
      ],
      [
        {
          "type": "str",
          "value": "HUE"
        }
      ],
      [
        {
          "type": "str",
          "value": "BET"
        }
      ],
      [
        {
          "type": "str",
          "value": "ROD"
        }
      ],
      [
        {
          "type": "str",
          "value": "NAC"
        }
      ],
      [
        {
          "type": "str",
          "value": "BAR"
        }
      ],
      [
        {
          "type": "str",
          "value": "SAM"
        }
      ],
      [
        {
          "type": "str",
          "value": "WAA"
        }
      ],
      [
        {
          "type": "str",
          "value": "EMP"
        }
      ],
      [
        {
          "type": "str",
          "value": "HAA"
        }
      ]
    ],
    "result_hash": "sha256:c276a97fdd20778fc5abe17922ef7d3f2e6eb13f75247763335e148959bb3166"
  },
  "planner_statistics": {
    "team": {
      "last_analyze": null,
      "last_autoanalyze": "2026-09-08 05:19:51.569042+00",
      "n_mod_since_analyze": 0
    },
    "team_attributes": {
      "last_analyze": null,
      "last_autoanalyze": "2026-09-08 05:19:47.099723+00",
      "n_mod_since_analyze": 0
    }
  },
  "shuffle": {
    "seed": "1",
    "row_limit": 300000,
    "tables": [
      "team",
      "team_attributes"
    ],
    "tables_not_shuffled": [],
    "tables_skipped_for_size": {
      "laptimes": 400524,
      "legalities": 427907,
      "posthistory": 303155,
      "trans": 1056320,
      "yearmonth": 383282
    },
    "tables_not_reached_by_a_copy": {}
  },
  "shuffled_copies": {
    "run": true,
    "verdict": "equal",
    "differs": false,
    "result_hash": "sha256:b2c3cded4853bdba02e359e8e4e7e643f79d4d1ff6506818831365f3aaf2cb7d",
    "result": {
      "columns": [
        {
          "name": "team_short_name",
          "declared_type": "text"
        }
      ],
      "row_count": 43,
      "truncated": false,
      "rows_shown": 10,
      "rows": [
        [
          {
            "type": "str",
            "value": "ARK"
          }
        ],
        [
          {
            "type": "str",
            "value": "ARL"
          }
        ],
        [
          {
            "type": "str",
            "value": "ARS"
          }
        ],
        [
          {
            "type": "str",
            "value": "BAR"
          }
        ],
        [
          {
            "type": "str",
            "value": "BET"
          }
        ],
        [
          {
            "type": "str",
            "value": "BMU"
          }
        ],
        [
          {
            "type": "str",
            "value": "BOL"
          }
        ],
        [
          {
            "type": "str",
            "value": "BRE"
          }
        ],
        [
          {
            "type": "str",
            "value": "CAG"
          }
        ],
        [
          {
            "type": "str",
            "value": "CAT"
          }
        ]
      ],
      "result_hash": "sha256:b2c3cded4853bdba02e359e8e4e7e643f79d4d1ff6506818831365f3aaf2cb7d"
    }
  },
  "plan_variant": {
    "run": false,
    "reason": "the plan variant was not asked for"
  }
}

The evidence records

gold: evidence-gold.json

SELECT DISTINCT t1.team_short_name FROM Team AS t1 INNER JOIN Team_Attributes AS t2 ON t1.team_api_id = t2.team_api_id WHERE t2.chanceCreationPassingClass = 'Safe'
statement read from
data/questions/mini_dev_postgresql.json
digest
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
origin
https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json, 2024-06-19

result_hash sha256:c276a97fdd20778fc5abe17922ef7d3f2e6eb13f75247763335e148959bb3166 recomputed from this JSON: match

record_hash sha256:700fa36bb36031da158b94afff40239101484b53fe862a1d8830091d9aefdcfe recomputed from this JSON: match

the result this record holds, 43 rows

from evidence-gold.json, 43 rows

team_short_nametext
WHU
HUE
BET
ROD
NAC
BAR
SAM
WAA
EMP
HAA
UTR
UDI
NAP
ARL
BMU
WIS
ARK
COR
REG
STK
SPA
PSV
FRE
CAG
ZAG
CAT
BRE
MCI
DUF
SIE
SAS
GEN
FRO
PAL
ARS
LOR
WII
LOK
BOL
LIV
GRF
COT
HER
what ran, and where
run
audit-34520748-ce64-4674-9cd7-6d6fac5b9805
executed at
2026-09-08T05:23:36.264626+00:00
data as of
2026-09-08T05:23:23.924963+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-SET
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
43 rows
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
search_path
"$user", public
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:9569693beaecfb53a614c6a17fa15de9183779d44aa279a4c366cb75b289dc45
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.team
299
rows in public.team_attributes
1458

second: evidence-second.json

SELECT Team.team_short_name
FROM Team
JOIN Team_Attributes ON Team.team_api_id = Team_Attributes.team_api_id
WHERE Team_Attributes.chanceCreationPassingClass = 'Safe'
statement read from
data/preds-pg/predict_mini_dev_gpt-4_postgresql.json
digest
sha256:cd39466740516d2d4672e8421010edab4816ff9875781f44df100840b41eaa49
origin
https://raw.githubusercontent.com/bird-bench/mini_dev/b3d4bcbbae9a96934ad812551eb400c7a3b23c12/llm/exp_result/sql_output_kg/predict_mini_dev_gpt-4_postgresql.json, 2024-06-19

result_hash sha256:00e1ae76585ce49c99428a6eae93166d84cbf7ad4afa7ecbc7c6747eea11e5ee recomputed from this JSON: match

record_hash sha256:0f5b31327e88f21edfa22751176eab3e6e84c353a24e83109da4eb4243e0beda recomputed from this JSON: match

the result this record holds, 56 rows

from evidence-second.json, 56 rows

team_short_nametext
HAA
ARK
ARL
ARS
ARS
ARS
BAR
BMU
BOL
BOL
BOL
BRE
CAG
CAG
CAT
COR
DUF
EMP
COT
UTR
FRE
FRO
GEN
GEN
GEN
GRF
HER
LIV
LIV
LOK
LOR
MCI
NAC
NAC
PAL
PSV
BET
HUE
REG
REG
SIE
ROD
SAM
SAS
SPA
NAP
STK
UDI
WAA
WHU
WHU
WII
WIS
WIS
WIS
ZAG
what ran, and where
run
audit-34520748-ce64-4674-9cd7-6d6fac5b9805
executed at
2026-09-08T05:23:36.268500+00:00
data as of
2026-09-08T05:23:23.924963+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-SET
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
56 rows
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
search_path
"$user", public
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:9569693beaecfb53a614c6a17fa15de9183779d44aa279a4c366cb75b289dc45
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.team
299
rows in public.team_attributes
1458

Running these again

gold

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

second

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

This question's run

run
audit-34520748-ce64-4674-9cd7-6d6fac5b9805
server
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
question set
mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
replay rule
R-SET

the run this question belongs to

The JSON this page was rendered from