rema / case files / case 01

my benchmark was grading its own homework

Diar Shakimov as a kid, at a judo tournamentDiar Shakimov6 min read

i grew the benchmark from 26 cases to 77 in a day. then found that the gold labels for 51 of them had been picked by the retriever the benchmark grades.

early august i audited my own bug-fixing agent project and gave its evaluation a 2 out of 10. there was a knowledge base behind the agent, and the claim "the knowledge base makes it better" couldn't be proven wrong. nothing measured which part helped, and every scorer change had been tuned on the same small set of tasks, with nothing held out.

the idea: a gold document for every solved bug

so i built the harness before building more features. every solved bug becomes a case file: the ticket, the reasoning and the diff. each case gets a gold document, a page that existed before the fix and that the retriever should have found. the eval checks how often it finds it. the retriever combines a keyword ranking with embeddings, fused by rank.

the obvious gold, the notes written after the fix, can't be used, because they'd leak the answer to the case they came from. so gold has to be a how-to or reference page that predates the fix, and someone has to go find it for each case. that labelling was the blocker. on august 14 there were 26 cases.

what didn't work: tripling the corpus in a day

so i wrote a pipeline to do it. it searched the ticket tracker with 102 keywords, and about 90% of the first results were wasted on two queues until i excluded them on the server side. it checked each candidate's change, author and ticket reference against the real fix and staged 90. then it picked a gold document for each and wrote the case file.

cases
case folders, august 14 to 1526 to 77
complete cases75
with a gold document72
gold picked by the pipeline51
gold labelled without the retriever21
staged, never written39
Table 1. the corpus on august 15. the 39 staged cases were never written, so they couldn't count toward anything.

why it was hard to see: the scores looked great

where each of the 77 cases got its gold label77 cases
Figure 1. the corpus on august 15, by who picked each case's answer key. 51 of 72 labelled cases, 71%, were graded against an answer the graded retriever had chosen. the hatched 5 are the cases with no gold label or an incomplete file (77 minus 72).

then i read how the pipeline picked gold. it called the retriever, took the top 5 results and kept one that matched by keywords and by embedding. that's the same retriever the benchmark grades. its first version even suggested another ticket's snapshot as gold, which needed an exclude list. so on 51 of the new cases, the score was mostly the retriever agreeing with itself. those 51 scored a lot higher than the independently labelled ones, and a blended average would have looked great and meant very little.

hamel husain's your ai product needs evals says human-labelled examples are what you use both to measure the system and to validate automated evaluation. this is what happens when you skip that step for speed: the automated part ends up validating itself.

the case: keeping 51 cases that can't count

the fix was structural. every case now records whether its gold label came from retrieval, and the summary prints the two groups separately, so no code path prints the blended number. the 51 are still real solved bugs, so i kept them. they just don't count as evidence until someone labels them without the retriever's help.

in september i re-ran the eval on a fresh checkout, and the headline number on the independent cases didn't reproduce. until i know why, there's no rate anywhere in this post.

"independent" is softer than it sounds, too. the 21 were labelled by hand or with an llm, and my own files disagree on the count: 21 in the write-up, 24 in the code comments.

if you hit this

for every gold label, record who or what picked it. if the system under test picked any of them, report those cases as their own group and never blend them into the headline.

what this changed in rema

a machine-picked answer key is never accuracy data. rema's eval builder only counts generated gold once independent raters have labelled it without seeing the machine's guess, and reports everything else separately.

Sources

Revised .

read more case files