The whole of this run is a release asset
This site holds 2 of the 13 question directories this run wrote; the archive holds every one:
- archive
- minidev-sqlite-gpt-4-turbo.tar.gz
- size
- 12,234,405 bytes
- sha256
- 4f44c336f425afe237d3fc1cb12c528c960d4f4a373a5f55d8c580eb73af4c63
What it counted
| what | how many |
|---|---|
| questions audited | 40 |
| EQUAL | 7 |
| NOT_EQUAL | 13 |
| ERROR | 20 |
| probes fired | 0 |
the probes
| probe | fired on |
|---|---|
| ordering-over-numeric-text | 0 |
| arbitrary-cut | 0 |
| not-a-function-of-the-data | 0 |
| float-aggregate-order | 0 |
| direction-against-question | 0 |
credited by BIRD's own check and NOT_EQUAL here
| what | how many |
|---|---|
| credited by BIRD and NOT_EQUAL here | 3 |
| multiplicity | 3 |
| type | 0 |
| order | 0 |
| truncation | 0 |
| other | 0 |
| the test-suite check answered 1 | 0 |
| the test-suite check answered 0 | 3 |
What the run was made of
- run
- audit-6d8d0e18-7d46-4a00-88ab-8c209953260e
- summary format
- attestql/audit/summary/2
- question file
- data/questions/mini_dev_sqlite.json sha256:4ba5fa8de55856222f484d380d2ba872b380bf79d825de70478e2120cb0fc43b | https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_sqlite.json, 2024-06-19
- questions in the file
- 500
- prediction file
- data/preds-sqlite/predict_mini_dev_gpt-4-turbo_sqlite.json sha256:ef9fcb0470d9710c8a9674fd35874ec0986925e27a4d5d9ccece8bff606fb5df | https://raw.githubusercontent.com/bird-bench/mini_dev/b3d4bcbbae9a96934ad812551eb400c7a3b23c12/llm/exp_result/sql_output_kg/predict_mini_dev_gpt-4-turbo_sqlite.json, 2024-06-19
- statements read
- 500
- keyed by
- position
- data file
- data/minidev/dev_databases/toxicology/toxicology.sqlite sha256:35ef27ae6bdfda530e125ed369666ac0abb8bc8c0bcc0ad09407f547ecb61a93 | minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f), member minidev/MINIDEV/dev_databases, toxicology/toxicology.sqlite, 2024-06-13
- server
- SQLite 3.53.4 | file=/private/tmp/attestql-runs/data/minidev/dev_databases/toxicology/toxicology.sqlite | size=2678784
- database role
- file
- engine
- sqlite
- parser
- validator audit:sqlglot-sqlite-parse, sqlglot 30.18.0, dialect sqlite
- serialization
- attestql/audit/2
- statement timeout
- 30 s
- fixture digest depth
- counts
- schema digest
- sha256:e31e13a0dd8834a30b867b1b7f46a0da62573bc5318da4230b24d801f34a519a
- data as of
- 2026-09-08T05:02:34.692321+00:00 the instant the run started
- shuffled copies
- prepared
- shuffle seed
- 1
- experimental probe
- off
The session the run was made in
what the engine reported, 9 settings
- sqlite_version
- 3.53.4
- encoding
- UTF-8
- reverse_unordered_selects
- 0
- query_only
- 1
- journal_mode
- delete
- data_version
- 1
- compile_options
- ATOMIC_INTRINSICS=1,COMPILER=clang-21.0.0,DEFAULT_AUTOVACUUM,DEFAULT_CACHE_SIZE=-2000,DEFAULT_FILE_FORMAT=4,DEFAULT_JOURNAL_SIZE_LIMIT=-1,DEFAULT_MMAP_SIZE=0,DEFAULT_PAGE_SIZE=4096,DEFAULT_PCACHE_INITSZ=20,DEFAULT_RECURSIVE_TRIGGERS,DEFAULT_SECTOR_SIZE=4096,DEFAULT_SYNCHRONOUS=2,DEFAULT_WAL_AUTOCHECKPOINT=1000,DEFAULT_WAL_SYNCHRONOUS=2,DEFAULT_WORKER_THREADS=0,DIRECT_OVERFLOW_READ,ENABLE_API_ARMOR,ENABLE_COLUMN_METADATA,ENABLE_DBSTAT_VTAB,ENABLE_FTS3,ENABLE_FTS3_PARENTHESIS,ENABLE_FTS5,ENABLE_GEOPOLY,ENABLE_MATH_FUNCTIONS,ENABLE_MEMORY_MANAGEMENT,ENABLE_PERCENTILE,ENABLE_PREUPDATE_HOOK,ENABLE_RTREE,ENABLE_SESSION,ENABLE_STAT4,ENABLE_UNLOCK_NOTIFY,MALLOC_SOFT_LIMIT=1024,MAX_ATTACHED=10,MAX_COLUMN=2000,MAX_COMPOUND_SELECT=500,MAX_DEFAULT_PAGE_SIZE=8192,MAX_EXPR_DEPTH=1000,MAX_FUNCTION_ARG=1000,MAX_LENGTH=1000000000,MAX_LIKE_PATTERN_LENGTH=50000,MAX_MMAP_SIZE=0x7fff0000,MAX_PAGE_COUNT=0xfffffffe,MAX_PAGE_SIZE=65536,MAX_SQL_LENGTH=1000000000,MAX_TRIGGER_DEPTH=1000,MAX_VARIABLE_NUMBER=250000,MAX_VDBE_OP=250000000,MAX_WORKER_THREADS=8,MUTEX_PTHREADS,SYSTEM_MALLOC,TEMP_STORE=1,THREADSAFE=1,USE_URI
- collation_list
- RTRIM,NOCASE,BINARY
- case_sensitive_like
- 0
What the run states about itself
- ids the question file states twice
- 137, 138
- prediction positions not compared
- 487, 488
The questions
| question | rule | verdict | class | probes | what was asked |
|---|---|---|---|---|---|
| q206 | R-SET | NOT_EQUAL | multiplicity | What elements are in the TR004_8_9 bond atoms? | |
| q268 | R-SET | NOT_EQUAL | multiplicity | What are the elements for bond id TR001_10_11? | |
| q200 | ERROR | prediction: statement: the text does not parse: Error tokenizing '.molecule_id = b.molecule_id WHERE b.bond_type = ' | |||
| q207 | ERROR | prediction: statement: the statement is a COLUMN and not a SELECT | |||
| q212 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'OUP BY element ORDER BY COUNT(element) ASC LIMIT ' | |||
| q215 | ERROR | prediction: statement: the statement is a COLUMN and not a SELECT | |||
| q218 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'le_id AND a.element = 'f' WHERE m.label = '+' | |||
| q219 | ERROR | prediction: statement: the text does not parse: Error tokenizing '0 / (SELECT COUNT(*) FROM bond WHERE bond_type = ' | |||
| q220 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'lecule_id = 'TR000' ORDER BY element ASC LIMIT ' | |||
| q227 | ERROR | prediction: statement: the text does not parse: Error tokenizing '_id)), 3) AS percentage_carcinogenic FROM molecul' | |||
| q231 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'ROUP BY bond_type ORDER BY bond_count DESC LIMIT ' | |||
| q232 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'D b.bond_type = '-' ORDER BY m.molecule_id LIMIT ' | |||
| q240 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'CT element FROM atom WHERE molecule_id = 'TR004' | |||
| q242 | ERROR | prediction: statement: the text does not parse: Error tokenizing '_id, 7, 2) BETWEEN '21' AND '25' AND m.label = '+' | |||
| q244 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'UP BY m.molecule_id ORDER BY COUNT(*) DESC LIMIT ' | |||
| q247 | ERROR | prediction: statement: the statement is a COLUMN and not a SELECT | |||
| q248 | ERROR | prediction: statement: the text does not parse: Error tokenizing ' WHERE b.molecule_id = 'TR041' AND b.bond_type = ' | |||
| q253 | ERROR | prediction: statement: the text does not parse: Error tokenizing 'ed.bond_id = bond.bond_id WHERE bond.bond_type = ' | |||
| q255 | ERROR | prediction: statement: the statement is a COLUMN and not a SELECT | |||
| q260 | ERROR | prediction: statement: the text does not parse: Error tokenizing ' c ON a.atom_id = c.atom_id WHERE b.bond_type = ' | |||
| q281 | ERROR | prediction: statement: the statement is a COLUMN and not a SELECT | |||
| q282 | ERROR | prediction: statement: the text does not parse: Error tokenizing '.molecule_id = 'TR006' GROUP BY m.molecule_i' |
A question whose statement could not be run wrote no directory, so it is a row here and has no page: the summary holds the side that stopped, the step it stopped at and the engine's own message.
The same index, restricted
Each of these is a page of its own, so a filtered view has an address a reader can send. Nothing here is done by a script or by a query string.
The JSON this page was rendered from
- summary.json attestql/audit/summary/2