Does anyone else cave?

The Claude dictionary, pointed at every other coding agent this developer uses.

Everything on the main page measures Claude models inside Claude Code. This page is a parallel check: the same phrase dictionary, the same matchers, the same rules, run over the transcripts of the other agents on the same developer's machine — . Nothing here feeds back into the Claude numbers; it sits beside them.

The unit is the model, not the tool it ran in. Where one model ran in several harnesses its messages are pooled and the harness is shown as a badge. The same -message floor applies; models below it are listed, greyed, with their intervals.

Concession rate by model per assistant message · 95% CI · every harness

Read with care: these are different conversations, different tasks, and different tools — a Codex session and a Claude Code session are not the same kind of afternoon. Claude rows come from the main corpus; the others from one machine.

Is the dictionary hearing them? the validity check that has to come before the comparison

The dictionary was built by enumerating what Claude says when it folds. A model that concedes in other words would be undercounted. So for each non-Claude model, every assistant message carrying any broadly fold-shaped language — apologies, "you're right", "my mistake", "I misread", "point taken" and the rest — was checked against the dictionary: how many of those does it already count?

Read with care: coverage is measured on a broad candidate net, so a low figure means "audit this model's untracked forms before trusting its rate", not "the rate is wrong". Untracked forms are enumerated locally and go through the same precision audit as every Claude phrase before they are added.

What each model says when it folds top matched phrases · counts

Read with care: the Claude rows draw on tens of thousands of messages; the others on hundreds to a few thousand. Rare phrases here are anecdotes.

How the user talks to each tool frustration rate per user message, by harness

Read with care: this is one person's temper across tools, and tone is the strongest predictor of concession on the main page — so a tool that gets sworn at less should be expected to concede less, before any difference in the model is invoked.