in august the site had a live demo that drafted an audit finding with a model, from a serverless function. if anything went wrong, it served a cached example generated offline months earlier, so a hiccup mid-demo wouldn't break the page. it cost nothing and needed no new infrastructure, which was the point.
why it was hard to see: three causes, one answer
for a demo a handful of people would see, the fallback was a reasonable call. the problem was that three different things all ended in the same cached answer: no access code, no api key, or a failed model call. only the last one logged anything, so from the outside all three looked the same.
the call also didn't stream. the prompt asked the model to write its reasoning and then end with a fenced json block, and the judge agent writes a few sentences before its json.
what didn't work: the night, in order
| what i tried | what came back |
|---|---|
| aug 18, 22:27: ship live generation | the cached answer |
| add the key to the hosting environment | a plain 200 with no error line: the no-key path |
| two manual redeploys, four passes over build and runtime logs | cached |
| check the project, the production checkbox, whitespace, branch scope | cached |
| add credits, delete an old alias | the old alias starts returning not-found, one more thing to rule out |
| aug 20, after midnight: one test call | live, so i closed the issue on a single test |
| an unrelated redeploy, two other cases | cached again, issue reopened |
| aug 20, 16:36: raise the response budget from 768 to 2048 tokens | 4 of 4 live |
billing had the answer before the logs did. the key had been used once, 4 cents off the credit, so the key worked. the response budget was 768 tokens, and a full finding plus its json is longer than that. the model got cut off mid-object, the closing fence never came, parsing failed, and the route quietly served the cached example. my best guess for the one success after midnight is a short case that happened to fit.
what finally showed it was a narrow runtime-log window with the draft agent's own error, unparseable response, on 2 of 3 regenerate calls.
jim shore's fail fast argues a system should fail immediately and visibly, so a bug shows up where it happens instead of three steps later. for a demo people are watching, i still think failing soft was right. what i got wrong was failing silent. the fallback should have said why it fired, on the page and in the log.
the case: the fix, and what it didn't fix
the fix was raising the budget to 2048 in both agents and accepting a bare code fence. about 42 hours after the live path shipped, it worked, 4 calls out of 4. it still never checked why the model stopped, so a longer finding could hit the same wall. a few weeks later the demo was retired anyway: the page in early september, the api routes the week after.
if you hit this
make every fallback say why it fired, in the response and in the log. and check why the model stopped (the stop reason your provider returns) before you try to parse what it wrote.
what this changed in rema
a fallback has to say it's a fallback. rema's search always says whether its index answered, a broken index is never papered over with the built-in corpus, and when rema can't tell, it reports unknown instead of zero.
Sources
Revised .

