gold
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
this statement states no ordering of its own
sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84
R-SET EQUAL float-aggregate-order
student_club · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
What is the total cost of the pizzas for all the events?
the hint the set supplies: total cost of the pizzas refers to SUM(cost) where expense_description = 'Pizza'
NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.
The two statements name their result columns differently. A projection is compared by position and declared type, so these names did not decide the verdict. They are part of the canonical rendering each result is hashed under, so the two result hashes differ with them.
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
this statement states no ordering of its own
sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84
SELECT SUM(cost) AS "Total Cost of Pizzas"
FROM expense
WHERE expense_description = 'Pizza'
this statement states no ordering of its own
sha256:e2bd35a91d578f3631b999c0c80e0e03ef0a653bc2c14bfd0fa4e2ae670b2d64
The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.
from counterexample.json, 0 rows, up to 25 shown per side
| side | sumfloat4 |
|---|---|
| no rows |
from counterexample.json, 0 rows, up to 25 shown per side
| side | Total Cost of Pizzasfloat4 |
|---|---|
| no rows |
set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal
https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py
result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them
ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py
from counterexample.json, 1 row
| sumfloat4 |
|---|
| 600.11 |
from counterexample.json, 1 row
| Total Cost of Pizzasfloat4 |
|---|
| 600.11 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
this statement aggregates floating point numbers, so its last digits depend on the order the rows were summed in; the values agree to six significant digits
from smells.json, 1 row
| 600.11005 |
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84",
"baseline_result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "600.11"
}
]
],
"result_hash": "sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84"
},
"planner_statistics": {
"expense": {
"last_analyze": null,
"last_autoanalyze": null,
"n_mod_since_analyze": 32
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"expense"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:28a33fd0933dc8c2b7eb34dc1b147cee178aa7665999cce122d919d75a4750d2",
"result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "600.11005"
}
]
],
"result_hash": "sha256:28a33fd0933dc8c2b7eb34dc1b147cee178aa7665999cce122d919d75a4750d2"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
},
"float_cells": [
{
"row": 0,
"column": "sum",
"declared_type": "float4",
"baseline": "600.11",
"rerun": "600.11005"
}
],
"significant_digits": 6
}
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
result_hash sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84 recomputed from this JSON: match
record_hash sha256:1da744534c8a2ac180244f01de80a1f740e4c3fc6d7799aaab48f6f2217a71c8 recomputed from this JSON: match
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 600.11 |
SELECT SUM(cost) AS "Total Cost of Pizzas"
FROM expense
WHERE expense_description = 'Pizza'
result_hash sha256:e2bd35a91d578f3631b999c0c80e0e03ef0a653bc2c14bfd0fa4e2ae670b2d64 recomputed from this JSON: match
record_hash sha256:2d79b8a659c076ecfea8143d70382fb3006e31ae561ba2dfc654e66c5ab70a5b recomputed from this JSON: match
from evidence-second.json, 1 row
| Total Cost of Pizzasfloat4 |
|---|
| 600.11 |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET