Your team's intelligence stays yours, and it compounds.

Rema helps any model, Claude, GPT, Gemini, Kimi, GLM or Llama, find where to look, cite the source, remember the fix and catch what's out of date.

Answers come from what your team already wrote, with the source, and every session adds to what the next person starts from. Nobody has to get good at prompting first.

Claude Code v2.1.283

Opus 5.5 · Claude Max

…/ml-service

❯ our pyproject sets index-strategy = unsafe-any-match so torch comes from the pytorch index. is that still right, and is it safe?

Searched for 1 pattern, read 2 files (ctrl+o to expand)

⏺ No to both. The setting is outdated, and it's risky for this project in

particular.

It's out of date. unsafe-any-match is now just a deprecated alias for

unsafe-first-match. uv only still accepts it for backwards compatibility

[https://github.com/astral-sh/uv/issues/2775#issuecomment-2129535430]. uv now

has three values: first-index (the default), unsafe-first-match and

unsafe-best-match [uv-threads/docs/concepts/indexes.md].

✻ Cooked for 14s · done 9:36 PM

Real recorded Claude Code session, with Rema reading public uv threads and docs, captured 2026-09-26. Rema adds one row; the answer crosses out the stale setting itself and cites the maintainer's reply. Full transcript

The Rema film: an engineer types one sentence asking Rema to fix a public GitHub issue; the recorded Claude Code session runs sped up; Rema hands the ticket to a teammate with what was tried and ruled out, and their agent picks up from there; then the title, rema.
Two real Claude Code sessions on a public GBrain issue, sped up. Rema shows what it pulled, each with its source; when the first session runs out of road, it hands the ticket to a teammate with what was tried and ruled out, and their agent picks up from there.

One person's win becomes everyone's head start.

The fix is kept with the evidence it came from, so it reaches the next person who asks, and their coding agent, without anyone repeating it.

Exampleagent session, March 12

Retries happen in the worker, per the 2023 design doc.

Not any more. Retries moved to the job queue in March. The design doc is out of date.
A staff engineer, correcting an answer once
  1. A new hire, week oneanswer cites the March 12 correction

    Why don't webhook retries live in the worker?

  2. On call, 2amanswer cites the March 12 correction

    Webhook failures are piling up. Where do retries run?

  3. A coding agent, mid-fixanswer cites the March 12 correction

    Which service owns the retry window?

  4. A PM, writing the specanswer cites the March 12 correction

    Can we make retries back off for longer?

What broke, and what changed.

Rema started as research, a paper on AI that judges claims against the evidence behind them, and became a product one fix at a time. Four of those fixes, each with somewhere you can check it.

Research pipeline

The loader read a slice and called it the file

Then we checked the humans: expert reviewers pushed back on truncated verdicts about as often as on complete ones. They did not catch it, so the system has to say how much of a file it read.

How it broke, and the fix
Broke
The spreadsheet loader kept only the first rows and columns of a large workbook and said nothing, so some evaluation verdicts were built on files it had only partly read.
Fix
Count every chunk that goes in, give an incomplete ingest its own label instead of a normal verdict, and lift the caps so a large file is read whole.
Rema

The live demo kept serving a months-old cached answer

The runtime logs had shown the unparseable-response errors the whole time. Now a failed call says it failed instead of passing an old answer off as new.

How it broke, and the fix
Broke
Everyone blamed the API key. The real bug: the response budget was too small, so the model's JSON was cut off mid-object, parsing failed, and the route fell back to a cached example without saying so.
Fix
Raise the budget and accept a bare code fence. The fallback stopped being silent.
Rema

A timeout could come back looking like an answer

A rule you can see in the CLI today: an empty index answers health unknown instead of a zero, and a source without its token writes nothing instead of reporting an empty result.

How it broke, and the fix
Broke
In the old research pipeline, a failed call could surface as "insufficient evidence", which reads exactly like a careful verdict.
Fix
Written into the clean-room rebuild on day one: an infrastructure failure is a diagnostic attached to the run, never a label on the answer.
Eval harness

The benchmark was grading its own homework

The corpus grew without moving the headline. Only independently labeled cases count as evidence; the rest are real solved tickets that wait for an independent review.

How it broke, and the fix
Broke
Growing the bug-fixer's case corpus, the gold document for many new cases was suggested by the same retrieval function the benchmark grades. Blended into one number, those cases would have quietly inflated the score.
Fix
Every case now records whether its gold label came from retrieval, and the summary prints the two populations separately, so the blended number can no longer be printed at all.

It reads the tools your team already writes in.

Threads, pull requests, tickets and pages, where the decisions actually happened. Anything else, a tracker, analytics, a database, comes in as an export: CSV, JSON or Markdown, each row and section kept with its place. If your tool is missing, say which one; that is how the next connector gets picked.

Start from what your team already wrote.

Rema is in a closed pilot. Point it at what your team already wrote; there is no reorganizing first.