gold
SELECT SUM(spent) FROM budget WHERE category = 'Food'
this statement states no ordering of its own
sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0
R-SET GOLD-ONLY float-aggregate-order
student_club · mini_dev_pg-00000-of-00001 from https://huggingface.co/datasets/birdsql/bird_mini_dev/resolve/f65faf4ae3b638c1fa6df1d3370c8d92c8366301/data/mini_dev_pg-00000-of-00001.json (commit f65faf4a, downloaded 2026-09-08)
What is the total amount of money spent for food?
the hint the set supplies: total amount of money spent refers to SUM(spent); spent for food refers to category = 'Food'
This question was audited without a prediction beside it, so there is nothing to compare the gold with. The probes below read the gold alone.
SELECT SUM(spent) FROM budget WHERE category = 'Food'
this statement states no ordering of its own
sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 1234.7299 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
this statement aggregates floating point numbers, so its last digits depend on the order the rows were summed in; the values agree to six significant digits
from smells.json, 1 row
| 1234.73 |
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0",
"baseline_result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "1234.7299"
}
]
],
"result_hash": "sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0"
},
"planner_statistics": {
"budget": {
"last_analyze": null,
"last_autoanalyze": null,
"n_mod_since_analyze": 52
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"budget"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:815419e4e8923f72ebbe635f8e89724feaf1a5f55f118349f98175b130cb1136",
"result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "1234.73"
}
]
],
"result_hash": "sha256:815419e4e8923f72ebbe635f8e89724feaf1a5f55f118349f98175b130cb1136"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
},
"float_cells": [
{
"row": 0,
"column": "sum",
"declared_type": "float4",
"baseline": "1234.7299",
"rerun": "1234.73"
}
],
"significant_digits": 6
}
SELECT SUM(spent) FROM budget WHERE category = 'Food'
result_hash sha256:abaa01a4ef80cd27d53cb3b55b4ba1df448a0d110ec3b5e034976199556be8a0 recomputed from this JSON: match
record_hash sha256:fb67abfe2aa60742b165c8e3e48ea11c042b4581fccb15a4b642858d301f4426 recomputed from this JSON: match
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 1234.7299 |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET