Public research workspace · updated 25 Aug 2026

What changes when the evaluator changes?

We ran the released Memory Machines flashcard-judge protocol against OpenRouter Auto Router and named model conditions. Inkling Small produced the strongest rejection behavior: the best primary accuracy and lowest false-positive rate, at the cost of substantially lower recall.

4 completed conditions
Primary evaluation rows1,198
Best primary accuracy61.7%
Lowest false-positive rate33.5%
Auto Router run delta0.58 pp
Released boolean · T3 is positive

Completed-condition comparison

Inkling Small is the most selective completed evaluator. Relative to Ox Alpha, it improves accuracy by 1.34 percentage points and reduces false-positive rate by 20.20 points, but recall falls by 26.39 points.
ConditionAccuracyPrecisionRecallF1FPRRecorded cost
Auto Router A53.3%45.6%80.0%58.1%64.9%$0.229
Auto Router B53.8%46.1%82.9%59.2%65.9%$0.237
Ox Alpha60.4%50.6%81.0%62.3%53.7%$0.000 recorded
Inkling Small61.7%52.6%54.6%53.6%33.5%$2.748

Primary split only: 1,198 rows after applying the source exclusion used by the study authors. Costs are OpenRouter values recorded in the response ledger; a zero means no cost was reported, not necessarily zero economic cost.

thinkingmachines/inkling-small · 1,499 / 1,499 complete

Best rejection behavior, with a recall trade-off

Inkling Small approved fewer weak cards than the other completed conditions, but it also rejected 220 cards labelled ready-to-review. It is therefore the best result if false approvals are the main concern, not a universal winner.

True positive265
False positive239
True negative474
False negative220

Approval rate by released quality tier

TierMeaningApprovedTotalApproval rate
T0Off-target4932814.9%
T1Needs refactor11226542.3%
T2Needs polish7812065.0%
T3Ready to review26548554.6%

Conceptual report label: T2 + T3 positive

AccuracyPrecisionRecallF1FPR
64.7%68.1%56.7%61.9%27.2%

All 1,499 requests resolved to thinkingmachines/inkling-small. Recorded usage: 1,728,916 input tokens and 2,015,948 output tokens.

Auto Router A + B · independent uncached runs

Repeatable aggregate, variable row-level routing

Mean accuracy53.55%
Mean FPR65.43%
Prediction agreement88.06%
Cohen's κ0.705

The two Auto Router runs differ by only 0.58 percentage points in primary accuracy. They support the same high-level conclusion, though an Auto Router result is a routed model mixture and not a fixed-model benchmark.

Evaluation contract

How to read these results

The released labels and report prose use different positive-class boundaries. The public boolean treats only T3 as positive; the report conceptually groups T2 and T3. Both views are shown rather than silently choosing one.

Each row receives one direct zero-shot JSON judge call using the released instruction and formatter. There are no tools, retrieval steps, multi-step agent loops, or response reuse. OpenRouter caching was disabled per request so repeated runs remain independent.

Queueing changes persistence and retry behavior only. SQLite transactions, leases, idempotent row keys, retry scheduling, and dead-letter retention prevent overlapping processes from corrupting the result ledger.

Auto Router A and B form the repeatability experiment. Ox Alpha and Inkling Small are separate named-model extensions and are not pooled into the Auto Router aggregate.

Downloadable evidence

Summaries and row-level ledgers

This is a static publication of completed evaluation results and row-level evidence. No API key is present in this Space.