q736

R-ORD EQUAL arbitrary-cut not-a-function-of-the-data

superhero · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json

The question

Who is the dumbest superhero?

the hint the set supplies: the dumbest superhero refers to MIN(attribute_value) where attribute_name = 'Intelligence'

NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.

The statements

gold

SELECT T1.superhero_name FROM superhero AS T1 INNER JOIN hero_attribute AS T2 ON T1.id = T2.hero_id INNER JOIN attribute AS T3 ON T2.attribute_id = T3.id WHERE T3.attribute_name = 'Intelligence' ORDER BY T2.attribute_value NULLS FIRST LIMIT 1
t2.attribute_value
ascending, nulls first

sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e

second

SELECT s.superhero_name
FROM superhero s
JOIN hero_attribute ha ON s.id = ha.hero_id
JOIN attribute a ON ha.attribute_id = a.id
WHERE a.attribute_name = 'Intelligence'
ORDER BY ha.attribute_value ASC
LIMIT 1;
ha.attribute_value
ascending, nulls default

sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e

The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.

The rows they differ in

in gold, not in the prediction, 0 rows

from counterexample.json, 0 rows, up to 25 shown per side

side superhero_nametext
no rows

in the prediction, not in gold, 0 rows

from counterexample.json, 0 rows, up to 25 shown per side

side superhero_nametext
no rows

What the benchmark would have said

BIRD's own check: 1

set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal

https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py

gold_rows
1
second_rows
1
gold_distinct_rows
1
second_distinct_rows
1

the test-suite check: 1

result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them

ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py

order_matters
true
gold_rows
1
second_rows
1
gold_columns
1
second_columns
1

The results

gold, 1 row

from counterexample.json, 1 row

superhero_nametext
Ammo
second, 1 row

from counterexample.json, 1 row

superhero_nametext
Ammo

The probes

A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.

ordering-over-numeric-text not applicable

this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10

no ORDER BY key resolves to a text column

what it measured
{
  "heuristic": true,
  "reason": "no ORDER BY key resolves to a text column",
  "keys": [
    {
      "key": "t2.attribute_value",
      "column": "hero_attribute.attribute_value",
      "declared_type": "bigint",
      "not_applicable": "the column is not declared as text"
    }
  ]
}

arbitrary-cut fired

this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero

from smells.json, 3 rows

Ammo 35
Jack-Jack 35
Ando Masahashi 35
what it measured
{
  "heuristic": true,
  "cut": 1,
  "offset": 0,
  "distinct_kept": false,
  "unbounded_sql": "SELECT t1.superhero_name, t2.attribute_value AS attestql_ordering_key_0 FROM superhero t1 JOIN hero_attribute t2 ON t1.id = t2.hero_id JOIN attribute t3 ON t2.attribute_id = t3.id WHERE t3.attribute_name = 'Intelligence' ORDER BY t2.attribute_value NULLS FIRST",
  "unbounded_rows": 623,
  "projected_columns": [
    "superhero_name"
  ],
  "ordering_key_columns": [
    "attestql_ordering_key_0"
  ],
  "ordering_keys": [
    {
      "key": "t2.attribute_value",
      "direction": "asc",
      "nulls": "first",
      "nulls_first_in_effect": true,
      "returned_rows_null_in_this_key": 0,
      "fires": false
    }
  ],
  "tied_at_the_cut": {
    "positions": [
      0,
      1,
      2
    ],
    "tied_rows": 3,
    "distinct_projected_answers": 3,
    "rows": [
      [
        {
          "type": "str",
          "value": "Ammo"
        }
      ],
      [
        {
          "type": "str",
          "value": "Jack-Jack"
        }
      ],
      [
        {
          "type": "str",
          "value": "Ando Masahashi"
        }
      ]
    ]
  },
  "case": "tie-at-the-cut"
}

not-a-function-of-the-data fired

rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data

from smells.json, 1 row

Jack-Jack
what it measured
{
  "heuristic": true,
  "rule": "R-ORD",
  "baseline_result_hash": "sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e",
  "baseline_result": {
    "columns": [
      {
        "name": "superhero_name",
        "declared_type": "text"
      }
    ],
    "row_count": 1,
    "truncated": false,
    "rows_shown": 1,
    "rows": [
      [
        {
          "type": "str",
          "value": "Ammo"
        }
      ]
    ],
    "result_hash": "sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e"
  },
  "planner_statistics": {
    "attribute": {
      "last_analyze": null,
      "last_autoanalyze": null,
      "n_mod_since_analyze": 6
    },
    "superhero": {
      "last_analyze": null,
      "last_autoanalyze": "2026-09-08 05:19:12.518866+00",
      "n_mod_since_analyze": 0
    },
    "hero_attribute": {
      "last_analyze": null,
      "last_autoanalyze": "2026-09-08 05:19:27.977203+00",
      "n_mod_since_analyze": 0
    }
  },
  "shuffle": {
    "seed": "1",
    "row_limit": 300000,
    "tables": [
      "attribute",
      "superhero",
      "hero_attribute"
    ],
    "tables_not_shuffled": [],
    "tables_skipped_for_size": {
      "laptimes": 400524,
      "legalities": 427907,
      "posthistory": 303155,
      "trans": 1056320,
      "yearmonth": 383282
    },
    "tables_not_reached_by_a_copy": {}
  },
  "shuffled_copies": {
    "run": true,
    "verdict": "not_equal",
    "differs": true,
    "result_hash": "sha256:e4a207112883b7b39b34cb7641405b3646153bf6487240f4d5b09afc08a78e3e",
    "result": {
      "columns": [
        {
          "name": "superhero_name",
          "declared_type": "text"
        }
      ],
      "row_count": 1,
      "truncated": false,
      "rows_shown": 1,
      "rows": [
        [
          {
            "type": "str",
            "value": "Jack-Jack"
          }
        ]
      ],
      "result_hash": "sha256:e4a207112883b7b39b34cb7641405b3646153bf6487240f4d5b09afc08a78e3e"
    }
  },
  "plan_variant": {
    "run": false,
    "reason": "the plan variant was not asked for"
  }
}

The evidence records

gold: evidence-gold.json

SELECT T1.superhero_name FROM superhero AS T1 INNER JOIN hero_attribute AS T2 ON T1.id = T2.hero_id INNER JOIN attribute AS T3 ON T2.attribute_id = T3.id WHERE T3.attribute_name = 'Intelligence' ORDER BY T2.attribute_value NULLS FIRST LIMIT 1
statement read from
data/questions/mini_dev_postgresql.json
digest
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
origin
https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json, 2024-06-19

result_hash sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e recomputed from this JSON: match

record_hash sha256:adfd3ebf6629fdc023db33914ea7c5bdda5024346261d3e80ae7c12866128faa recomputed from this JSON: match

the result this record holds, 1 row

from evidence-gold.json, 1 row

superhero_nametext
Ammo
what ran, and where
run
audit-715e02b9-6229-4aab-aa2e-7df62f68e7a4
executed at
2026-09-08T05:25:53.109336+00:00
data as of
2026-09-08T05:25:39.400422+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-ORD
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
1 row
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
search_path
"$user", public
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:4cc686540d6b0c0f1f775cd580e90e91fa664c968446deb705928be3775641fd
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.attribute
6
rows in public.hero_attribute
3738
rows in public.superhero
750

second: evidence-second.json

SELECT s.superhero_name
FROM superhero s
JOIN hero_attribute ha ON s.id = ha.hero_id
JOIN attribute a ON ha.attribute_id = a.id
WHERE a.attribute_name = 'Intelligence'
ORDER BY ha.attribute_value ASC
LIMIT 1;
statement read from
data/preds-pg/predict_mini_dev_phi-3-medium-128k-instruct-1_postgresql.json
digest
sha256:8defb4399ab3393200cfba510024df7d9cf2cff8c9c7df647e3309cf751ac82d
origin
https://raw.githubusercontent.com/bird-bench/mini_dev/b3d4bcbbae9a96934ad812551eb400c7a3b23c12/llm/exp_result/sql_output_kg/predict_mini_dev_phi-3-medium-128k-instruct-1_postgresql.json, 2024-06-19

result_hash sha256:c58fd60637a6677cc98abbe74ea8b745da3c6213017f5935773bb3e4a3d6177e recomputed from this JSON: match

record_hash sha256:bc5bd116a4edda3f8e997813ecb18b3c82f36568073e967dcbe58ed2ff42f1b4 recomputed from this JSON: match

the result this record holds, 1 row

from evidence-second.json, 1 row

superhero_nametext
Ammo
what ran, and where
run
audit-715e02b9-6229-4aab-aa2e-7df62f68e7a4
executed at
2026-09-08T05:25:53.112018+00:00
data as of
2026-09-08T05:25:39.400422+00:00
backend at checkout
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
backend that answered
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
database role
auditor
replay rule
R-ORD
question set version
sha256:d2731292f20b8d8569cd956dd747ffe1df13cd625076263e38ae9ebcef50b1ab
validator
audit:libpg_query-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
120000 ms
rows
1 row
the session it ran under
engine
postgresql
time_zone
Etc/UTC
date_style
ISO, MDY
interval_style
postgres
extra_float_digits
1
database_collation
en_US.utf8
work_mem
4096
hash_mem_multiplier
2

recorded beside them

statement_timeout
0
search_path
"$user", public
server_version
16.15 (Debian 16.15-1.pgdg13+2)
server_version_num
160015
transaction_read_only
on
max_parallel_workers_per_gather
2
server_encoding
UTF8
datlocprovider
c
daticulocale
datcollversion
2.41
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:4cc686540d6b0c0f1f775cd580e90e91fa664c968446deb705928be3775641fd
source file sha256
sha256:31b1da211849d24a57c9af7636da46a5b82fc8a3ca1542bb3ebd8775e9a31cec
rows in public.attribute
6
rows in public.hero_attribute
3738
rows in public.superhero
750

Running these again

gold

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD

second

re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD

This question's run

run
audit-715e02b9-6229-4aab-aa2e-7df62f68e7a4
server
PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird
question set
mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
replay rule
R-ORD

the run this question belongs to

The JSON this page was rendered from