gold
SELECT SUM(spent) FROM budget WHERE category = 'Food'
this statement states no ordering of its own
sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0
R-SET GOLD-ONLY float-aggregate-order
student_club · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
What is the total amount of money spent for food?
the hint the set supplies: total amount of money spent refers to SUM(spent); spent for food refers to category = 'Food'
This question was audited without a prediction beside it, so there is nothing to compare the gold with. The probes below read the gold alone.
SELECT SUM(spent) FROM budget WHERE category = 'Food'
this statement states no ordering of its own
sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 1234.7299 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
this statement aggregates floating point numbers, so its last digits depend on the order the rows were summed in; the values agree to six significant digits
from smells.json, 1 row
| 1234.73 |
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0",
"baseline_result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "1234.7299"
}
]
],
"result_hash": "sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0"
},
"planner_statistics": {
"budget": {
"last_analyze": null,
"last_autoanalyze": null,
"n_mod_since_analyze": 52
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"budget"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:815419e4e8923f72ebbe635f8e89724feaf1a5f55f118349f98175b130cb1136",
"result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "1234.73"
}
]
],
"result_hash": "sha256:815419e4e8923f72ebbe635f8e89724feaf1a5f55f118349f98175b130cb1136"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
},
"float_cells": [
{
"row": 0,
"column": "sum",
"declared_type": "float4",
"baseline": "1234.7299",
"rerun": "1234.73"
}
],
"significant_digits": 6
}
SELECT SUM(spent) FROM budget WHERE category = 'Food'
result_hash sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0 recomputed from this JSON: match
record_hash sha256:6bf2d39553953e08e739c35edb7acb78fdf701cbcd63f91c6f7b3a38989fa1eb recomputed from this JSON: match
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 1234.7299 |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET