rema / case files / case 03

i capped my own evidence at 500 rows. most auditors didn't notice.

Diar Shakimov as a kid, at a judo tournamentDiar Shakimov7 min read

a loader cap i had added in march kept 500 of 1,500 rows and 50 of 65 columns, and said nothing. 6 of 48 evaluation cells ran on truncated files, and the auditors' concern rate barely moved.

in september i went back to the IntelliAudit research pipeline, the one behind our paper, to look at one control, the asset inventory. the pipeline is four agents on one model: search, auditor, defender and judge. the evaluation runs 12 controls against 4 simulated organizations, 48 cells, and the evidence for this control is a big spreadsheet. the findings read like the system had seen all of it.

what didn't work: three wrong answers

it hadn't, and it took three wrong answers to see why. the roadmap blamed a 150-line cap that doesn't exist. the paper draft blamed the model for being lazy and bet that a newer model would fix it, and the rerun on the newer model produced a parsing error instead. a later commit blamed a 120-row cap, which lives in the survey app and never touched the pipeline.

why it was hard to see: the caps were stacked

the real one was mine. back in march i had given the spreadsheet loader a cap of 500 rows and 50 columns. it kept those and dropped the rest, no warning. and it wasn't the only cap. retrieval returned a bounded number of chunks, the planner only saw the first 50 evidence items, and the judge saw 200-character excerpts. the sufficiency check that should have asked for more never ran, because its counter counted queries.

the fix: counting what got in

sheetin the fileloaded
asset inventory1,500 x 65500 x 50
lifecycle events2,973 rows500 rows
Table 1. what the loader kept from the first organization's workbook. 4,473 rows and 15 columns never reached the pipeline. the other two affected organizations lost 400 and 600 rows.

the auditor agent ended up seeing 131 to 140 of 2,395 chunks, about 6%, and two sheets were never returned at all. lifting the caps showed how much was missing.

organizationcappedcap-free
first2,3957,239
second1,5042,104
fourth1,5032,403
Table 2. chunks per organization, with the caps and without them.

the fixes went in over a few days. attestation first, so every chunk that goes in is counted and the system can say how much of a file it read. then a gate, so an incomplete ingest gets its own label and withholds the verdict instead of guessing. that one mattered beyond spreadsheets: a judge timeout had been coming back as INSUFFICIENT_EVIDENCE, and a planner parse failure as PARTIAL, and both read exactly like real findings. then runs with the caps lifted.

those cost more than i expected. attesting the smallest organization took 235k input tokens and 233 seconds. the first organization's cap-free run hit credit balance too low 64 times and finished 5,960 of 7,239 chunks. the one fully attested cap-free run, 981 of 981 chunks, was overwritten by a second run that attested 0, and now survives only in my notes.

the case: did the auditors notice?

we had real auditor ratings on these findings from the survey, so i checked whether they pushed back more on the truncated ones.

findingsconcernsrate
built on truncated files5 of 2520.0%
built on complete files25 of 13818.1%
Table 3. auditor ratings that raised a concern, by whether the finding was built on a truncated file. a concern means rating factual correctness below agree.

some did notice. 2 of the 7 auditors who rated one truncated finding named the omission exactly, one of them writing that they could only see the hardware assets. but the aggregate barely moved, so across careful, experienced reviewers, a finding built on a third of the file mostly passed. it's a small sample, and it measures perceived quality, not ground truth.

hamel husain tells people starting on evals to examine as much data as possible and read the traces from every test case. i agree, and this is the case for reading one level lower. every trace here looked fine. the problem was in what the loader let into the trace.

so the system has to say it, every time, in the finding itself.

if you hit this

for every file you ingest, log the rows and columns in the file next to the rows and columns you loaded. if they differ, say so in the output a person reads, not only in the log.

what this changed in rema

a truncated ingest is recorded as truncated. rema's spreadsheet extractor reads the whole workbook, coverage is never called complete if any file was cut, and a finding says so when no population-wide conclusion is available.

Sources

read more case files