gold
SELECT T1.`First Date` FROM Patient AS T1 INNER JOIN Laboratory AS T2 ON T1.ID = T2.ID WHERE T2.LDH < 500 ORDER BY T2.LDH DESC LIMIT 1
- T2.LDH
- descending
sha256:ece0fe2f3652a31a52eb5e2bdb25c4342defd3d86824a20a4d20a6a30da5a052
R-ORD GOLD-ONLY arbitrary-cut not-a-function-of-the-data
thrombosis_prediction · dev_20251106-00000-of-00001 from https://huggingface.co/datasets/birdsql/bird_sql_dev_20251106/resolve/3c11fb193e5439b338e23677fa0aae11e8b85db9/data/dev_20251106-00000-of-00001.json (commit 3c11fb19, downloaded 2026-09-07)
For the patient with the highest lactate dehydrogenase in the normal range, when was his or her data first recorded?
the hint the set supplies: highest lactate dehydrogenase in the normal range refers to MAX(LDH < 500); when the data first recorded refers to MIN(First Date);
This question was audited without a prediction beside it, so there is nothing to compare the gold with. The probes below read the gold alone.
SELECT T1.`First Date` FROM Patient AS T1 INNER JOIN Laboratory AS T2 ON T1.ID = T2.ID WHERE T2.LDH < 500 ORDER BY T2.LDH DESC LIMIT 1
sha256:ece0fe2f3652a31a52eb5e2bdb25c4342defd3d86824a20a4d20a6a30da5a052
from evidence-gold.json, 1 row
| First DateTEXT |
|---|
| 1992-09-21 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
no ORDER BY key resolves to a text column
{
"heuristic": true,
"reason": "no ORDER BY key resolves to a text column",
"keys": [
{
"key": "T2.LDH",
"column": "Laboratory.LDH",
"declared_type": "INTEGER",
"not_applicable": "the column is not declared as text"
}
]
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
from smells.json, 5 rows
| 1992-09-21 | 499 |
| 1995-08-08 | 499 |
| 1989-12-02 | 499 |
| 1990-05-08 | 499 |
| 1993-01-25 | 499 |
{
"heuristic": true,
"cut": 1,
"offset": 0,
"distinct_kept": false,
"unbounded_sql": "SELECT T1.\"First Date\", T2.LDH AS attestql_ordering_key_0 FROM Patient AS T1 INNER JOIN Laboratory AS T2 ON T1.ID = T2.ID WHERE T2.LDH < 500 ORDER BY T2.LDH DESC",
"unbounded_rows": 10016,
"projected_columns": [
"First Date"
],
"ordering_key_columns": [
"attestql_ordering_key_0"
],
"ordering_keys": [
{
"key": "T2.LDH",
"direction": "desc",
"nulls": "last",
"nulls_first_in_effect": false,
"returned_rows_null_in_this_key": 0,
"fires": false
}
],
"tied_at_the_cut": {
"positions": [
0,
1,
2,
3,
4
],
"tied_rows": 5,
"distinct_projected_answers": 5,
"rows": [
[
{
"type": "str",
"value": "1992-09-21"
}
],
[
{
"type": "str",
"value": "1995-08-08"
}
],
[
{
"type": "str",
"value": "1989-12-02"
}
],
[
{
"type": "str",
"value": "1990-05-08"
}
],
[
{
"type": "str",
"value": "1993-01-25"
}
]
]
},
"case": "tie-at-the-cut"
}
rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data
from smells.json, 1 row
| 1993-01-25 |
{
"heuristic": true,
"rule": "R-ORD",
"baseline_result_hash": "sha256:ece0fe2f3652a31a52eb5e2bdb25c4342defd3d86824a20a4d20a6a30da5a052",
"baseline_result": {
"columns": [
{
"name": "First Date",
"declared_type": "TEXT"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "str",
"value": "1992-09-21"
}
]
],
"result_hash": "sha256:ece0fe2f3652a31a52eb5e2bdb25c4342defd3d86824a20a4d20a6a30da5a052"
},
"planner_statistics": {},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"Patient",
"Laboratory"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:53cfb437129e842acadea2d5f43f008089557bd110f891745efa16655c63f788",
"result": {
"columns": [
{
"name": "First Date",
"declared_type": "TEXT"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "str",
"value": "1993-01-25"
}
]
],
"result_hash": "sha256:53cfb437129e842acadea2d5f43f008089557bd110f891745efa16655c63f788"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
}
}
SELECT T1.`First Date` FROM Patient AS T1 INNER JOIN Laboratory AS T2 ON T1.ID = T2.ID WHERE T2.LDH < 500 ORDER BY T2.LDH DESC LIMIT 1
result_hash sha256:ece0fe2f3652a31a52eb5e2bdb25c4342defd3d86824a20a4d20a6a30da5a052 recomputed from this JSON: match
record_hash sha256:fbf34a75b3eb37614de09144826020c3898bb3b4ce084577c47a25ac73c450f2 recomputed from this JSON: match
from evidence-gold.json, 1 row
| First DateTEXT |
|---|
| 1992-09-21 |
re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-runs/data/dev/dev_databases/thrombosis_prediction/thrombosis_prediction.sqlite | size=7327744 under the session settings and over the data this record's fixture digest names, and compare the two results under R-ORD