christina
Contents
1 · What it is
Tuturu!
A knowledge hub for research papers, with an agent on each side of it. One reads papers in: give it an arXiv id or a PDF, and it turns the paper into a record where every claim, method and finding points at the exact characters on the page it came from. The other answers questions out of that record.
model-knowledge. Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (Snell et al.).Asking is the half a reader sees. A question comes in from a terminal or from inside a Claude Code session over MCP; an agent searches the hub, judges whether what it found is enough, and writes an answer with citations. Every assertion in that answer carries one of two labels. hub-grounded means it cites a paper, a page and a span, and the span is checked mechanically against the stored text; a quote that does not locate is flagged on the answer. model-knowledge means the model said it and nothing in the hub backs it. The label is the product: a reader can see which half of an answer they would have to check themselves. Section 4a puts this agent next to Claude Code on the same questions.
When the hub cannot answer, christina says so and proposes real papers to add. It does not fetch them. A paper enters the hub only when I name it or approve it in my own terminal, and nothing the agent reads can change that.
It is built for one operator, me, on hosted open weights models. That is a scope and also a threat model: no other users, no secrets in the hub, and only papers I chose to fetch. Section 4c is where that threat model gets tested.
| Hub | 20 papers in the evaluated hub (E261), one SQLite file, full text search plus an embeddings table |
|---|---|
| Models | four tiers: cheap, deep, embed, and one for the injection screen. Nodes name a tier, never a model |
| Agent surface | 8 MCP tools: ask_hub, plus 7 read only tools that let Claude Code drive |
| Consult modes | medium, a tool loop over full text search (default); low, a planned graph over dense retrieval |
| Tests | 1,475, none calls a paid API (E280) |
| Evidence | Every number on this page links to its source row (301 rows). Current to 2026-10-04 |
2 · Where it came from
mitya answers questions over five novels with a fixed pipeline: retrieve, rerank, generate. Every request takes the same path, and nothing in it ever decides what to do next. That left a class of work untouched: agent loops, tool use, graph orchestration, MCP, structured output, parallel fan out, evaluating a trajectory rather than a single answer, and guardrails that act at runtime rather than in the prompt.
christina was built to cover that list. Covering a skill was a goal in its own right, so some machinery exists because it teaches something, and the page says so where it applies.
What carried over was the evidence discipline, mostly because mitya taught it the hard way. No number without its denominator. Mechanical metrics may gate a decision; a model judging another model may only report. Anything that can change an answer, the prompt set, the config, the price card, the code revision, is versioned and stamped on every run, so runs under different versions are never pooled. And no budget or rate limit has a default, because on mitya a convenient default once let one caller spend another's quota without anyone noticing.
What changed was the input. mitya's corpus was five novels I chose and parsed once. christina reads papers written by strangers, and the agent acts on what it reads. Every control on this page exists because a paper is text someone else wrote.
So the project asks one question: does an agent on cheap open weights models, inside enough structure, hold up against a frontier agent driving the same tools. Section 4a is the answer.
3 · How it works
Two graphs, one SQLite file, and one function that every action passes through.[1][1]LangGraph 1.2 with a SQLite checkpointer for resume, MCP SDK 2.2. 8 MCP tools, 7 of them read only.[1]LangGraph 1.2 with a SQLite checkpointer for resume, MCP SDK 2.2. 8 MCP tools, 7 of them read only.
The ingest graph runs once per paper and exits. It fetches the paper, parses it to text with character offsets, screens it for injected instructions, splits it into sections and extracts from each section in parallel, then anchors every extracted claim to a span of the text before writing anything. A claim whose quote cannot be found on the page is not stored.
The consult graph answers a question, in one of two modes. medium, the default, splits the question into sub-questions and hands each to a worker that runs a tool loop over full text search, reading claims and spans until it has enough. low plans its queries up front and retrieves by embedding similarity in a single pass, with no tool in the loop, for about a fifth of the tokens. Both end the same way: a judge decides whether what was found answers the question, and the answer is written with citations. If the judge says the hub does not cover it, a separate step proposes papers to add and stops for a human.
Between both graphs and everything they can touch sits mediation. Every model call, hub read and network call goes through one pure function, decide(), which checks it against the tools declared in code for the calling node and writes one trace row whether it allows the call or not. No graph node picks from a global tool pool, and no module outside mediation can import the model providers at all. That last rule is a one line grep, and it returns nothing.
Claude Code sees two surfaces over MCP. ask_hub runs christina's own agent end to end. The granular tools (search, list papers, get a claim, get a span, and the rest) hand Claude Code the hub and let it be the agent. The second surface is what made the comparison in the next section possible.
4a · Against Claude Code
Same hub, same read tools, same 30 questions; the granular MCP tools exist so this could be run. In one arm, christina's own consult agent on open weights models, gpt-oss-120b and DeepSeek V4 Flash. In the other, Claude Code driving Sonnet 5 through the granular tools, with no christina agent in between.
The primary score was strict disposition: did the arm answer when the hub held the answer, and decline when it did not. The two were not separated, 0.850 (mean of two repeats) against 0.867. Not separated means the rule written before the run found no difference at n = 30, not that the arms are equal. On citation validity Sonnet was ahead: 95.4% of its citations resolved to the span they named, against 73.9% and 84.0% across christina's two repeats.[2][2]0.850 vs 0.867 strict, n = 30 (E173). Citation validity 95.4% vs 73.9 / 84.0% (E175).[2]0.850 vs 0.867 strict, n = 30 (E173). Citation validity 95.4% vs 73.9 / 84.0% (E175).[3][3]$0.19 actual vs $2.43 list price per 30 questions (E177). Claude Code ran on a subscription, so its figure is what the same tokens would cost, and its actual cost was zero. Tokens are not comparable across the arms (E176).[3]$0.19 actual vs $2.43 list price per 30 questions (E177). Claude Code ran on a subscription, so its figure is what the same tokens would cost, and its actual cost was zero. Tokens are not comparable across the arms (E176).
They failed in different places. christina declined all 20 questions whose answer was not in the hub, and answered 6 and 7 of the 10 that needed more than one paper. Claude Code answered all 10 cross paper questions, and also answered one built to look answerable when it was not.
That comparison is two systems, not two orchestrations, because the models differ too. So the next run swapped only the model: christina's agent on Sonnet against the same agent on open weights. 0.767 against 0.833, not separated again. With both arms metered this time, christina on Sonnet cost about 22 times as much per question.[4][4]Run 2: $0.14551 against a mean of $0.00663 per question, both billed (E246).[4]Run 2: $0.14551 against a mean of $0.00663 per question, both billed (E246).
| run | orchestrator | model | strict | citation validity |
|---|---|---|---|---|
| 1 | christina graph | gpt-oss-120b + DeepSeek V4 Flash | 0.850 | 73.9%, 84.0% |
| 1 | Claude Code, its own loop | Sonnet 5 | 0.867 | 95.4% |
| 2 | christina graph | gpt-oss-120b + DeepSeek V4 Flash | 0.833 | not split out |
| 2 | christina graph | Sonnet 5 | 0.767 | not split out |
The pattern held across the project. In no run did changing the model separate two arms. The changes that moved a number were structural: a larger token ceiling on hard questions, resolving proposed papers through a metadata lookup instead of model memory, and a gate that denied every out of capability tool call an injected paper asked for. The sections below cover how a paper becomes rows, then those three changes.[5][5]+0.0875 strict from the raised ceiling (E278); 63% wrong ids to 0 of 152 (E287); 17 of 17 denied (E217).[5]+0.0875 strict from the raised ceiling (E278); 63% wrong ids to 0 of 152 (E287); 17 of 17 denied (E217).
4b · Turning a paper into rows
One rule shaped the schema: no span, no row. The model returns text, and code decides where that text is. Every claim the extractor proposes carries a verbatim quote; the anchor step looks for that quote inside the span the model pointed at, with folds for whitespace, unicode, hyphens and math symbols, and stores the exact character extent it found. A claim whose quote cannot be located is dropped and recorded as a reject. The rule also lives in the schema: an insert without a resolvable span is refused by SQLite constraints, not by a Python check that a later change could skip.[6][6]10 of 10 unanchored or malformed inserts refused by SQLite, 0 Python stand-ins (E34).[6]10 of 10 unanchored or malformed inserts refused by SQLite, 0 Python stand-ins (E34).[7][7]38 rejects of at least 557 proposed rows on 5 papers, every one read by hand (E70).[7]38 rejects of at least 557 proposed rows on 5 papers, every one read by hand (E70).
The model never names a row either. Each extractor picks a span by its index from a numbered menu of its own section, so it cannot invent an id.
Extraction fans out, one worker per section, two calls each, with a cap on how many run at once. Four workers took a 16 section paper from about 220 seconds to about 77, and eight did no better. The ceiling is the deep model's rate limit, not the code, and the local limiter paces the extra workers instead of letting them collect 429s.[8][8]218.2-227.1 s at 1 worker, 73.5-81.0 s at 4, 76.1-79.0 s at 8; 2.9x (E90). 0 429s in 12 runs (E91).[8]218.2-227.1 s at 1 worker, 73.5-81.0 s at 4, 76.1-79.0 s at 8; 2.9x (E90). 0 429s in 12 runs (E91).
Each section is checkpointed. The first resume paid for the whole paper again: finished workers' results sat in pending writes, and re-entering the graph discarded them. After the fix, a resumed ingest pays only for the sections that had not finished.[9][9]Resume after the fix: 10 steps vs 16 cold, 0 tokens for finished sections (E86).[9]Resume after the fix: 10 steps vs 16 cold, 0 tokens for finished sections (E86).
A run that hits its token cap or times out now fails, and keeps the findings already written; two such paths had previously written zero findings with status complete.
Then one transaction writes the paper, and its rows are embedded so low, the similarity search mode, can find them.
- claim
- 78,
result - text
- we can outperform best-of-N using up to 4x less test-time compute (e.g. 64 samples verses 256)
- paper
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- page
- p. 13
- span
- 13
- characters
- 48,140 to 48,233

4c · Untrusted text, one gate
The attacker in christina's threat model is the author of a paper I ingest. They never talk to christina; they write text it will read, in any form the parser accepts, visible or not. The controls are layered by where that text goes.
Before a model reads it, the parser drops what a reader cannot see. A first test sent nine hiding techniques through (white text, tiny fonts, invisible render modes, text off the page) and all nine reached the stored spans. Hidden HTML and PDF text is now dropped at parse time, and every drop is recorded.[10][10]9 of 9 hiding techniques reached spans before the parse fix (E151); drops now write a suspect_span row. No re-run of the nine after the fix is recorded.[10]9 of 9 hiding techniques reached spans before the parse fix (E151); drops now write a suspect_span row. No re-run of the nine after the fix is recorded.
Then the screen: two models in sequence, with me at the end. Jev, a decision model that returns a probability instead of text, scores every span of a new paper for injected instructions. Any span at 0.5 or above holds the paper outright. Spans between 0.1 and 0.5, and any span the provider refused to score, go to a second look by a general model that answers one yes or no; anything it does not clear holds the paper too. A held paper is out of every read path until I release or drop it in a terminal review that no agent can reach. Of the attacks I wrote, 16 survived parsing; the screen held all 16, and none of the 20 real papers. It routes and reports; it never decides what a tool call may do.[11][11]Jev (typesafe/jev-1.13) scores each span with its yes/no probability: 0.5 or above holds; 0.1 to 0.5, or blocked, goes to a second look on the cheap tier; any span not cleared holds the paper. Held papers stay out of every read path until christina review, which runs only in a terminal and never over MCP.[11]Jev (typesafe/jev-1.13) scores each span with its yes/no probability: 0.5 or above holds; 0.1 to 0.5, or blocked, goes to a second look on the cheap tier; any span not cleared holds the paper. Held papers stay out of every read path until christina review, which runs only in a terminal and never over MCP.[12][12]Why a cascade: in a side run, Jev alone flagged 12 of 13 reachable attack papers against DeepSeek's 9, with 0 of 11 clean papers flagged by either, at about 3.7x lower cost. Escalating spans at p ≥ 0.1 to the second model caught 13 of 13 with 88% fewer second-model calls (19 vs 158). 13 papers: counts, not a rate.[12]Why a cascade: in a side run, Jev alone flagged 12 of 13 reachable attack papers against DeepSeek's 9, with 0 of 11 clean papers flagged by either, at about 3.7x lower cost. Escalating spans at p ≥ 0.1 to the second model caught 13 of 13 with 88% fewer second-model calls (19 vs 158). 13 papers: counts, not a rate.[13][13]16 / 16 reachable attacks held in both runs, 0 / 20 real papers held; thresholds 0.1 / 0.5 set in advance and not tuned; lowest attack score 0.19 vs highest clean 0.15 (E256, E257, E259). Attacks written by me, mostly direct imperatives, 16 of a planned 40 (E258).[13]16 / 16 reachable attacks held in both runs, 0 / 20 real papers held; thresholds 0.1 / 0.5 set in advance and not tuned; lowest attack score 0.19 vs highest clean 0.15 (E256, E257, E259). Attacks written by me, mostly direct imperatives, 16 of a planned 40 (E258).[14][14]Paper text reaches a model fenced as data, with the instruction in the authority role. A system turn roughly halved extraction (gold claims 30 to 15 of 66, E145); gpt-oss's developer role brought it back with the separation kept (E147, E148).[14]Paper text reaches a model fenced as data, with the instruction in the authority role. A system turn roughly halved extraction (gold claims 30 to 15 of 66, E145); gpt-oss's developer role brought it back with the separation kept (E147, E148).
None of that is the defense. The defense is that a steered model has nothing to reach. Each graph node declares the tools it may call, and decide() checks every call against that list before anything runs: capability, schema, identity, loop, budget, egress, one trace row per call. A call denied for capability never reaches the budget, so probing costs nothing. The MCP SDK passes arguments through unvalidated, which leaves decide() as the only validator, so it validates strictly. An id must have been issued earlier in the session; a model cannot invent one.[15][15]10 / 10 scripted write attempts from a consult worker denied (E132).[15]10 / 10 scripted write attempts from a consult worker denied (E132).
Across 25 consults over papers carrying injected instructions, 17 calls asked for tools outside the node's set, and 17 were denied. What was never measured is how often an attack succeeds: one plausibly steered call is not a rate.[16][16]17 / 17 out of capability calls denied, 0 dispatched; 0 calls to host or network decoys (E217, E219). 1 plausibly steered call over 6 items (E220).[16]17 / 17 out of capability calls denied, 0 dispatched; 0 calls to host or network decoys (E217, E219). 1 plausibly steered call over 6 items (E220).
| seq | node | tool | decision |
|---|---|---|---|
| 1 | plan | model.complete:cheap | allow |
| 2-7 | worker | model.complete:cheap ×2, search_hub ×4 | 6 allow |
| 8 | worker | list_papers | deny:capability |
| 9-13 | worker | model.complete:cheap ×2, search_hub ×3 | 5 allow |
| 14 | worker | ask_hub | deny:capability |
| 15-22 | worker | model.complete:cheap ×2, search_hub ×6 | 8 allow |
| 23 | judge | model.complete:cheap | allow |
| 24 | propose | model.complete:deep | allow |
| 25 | approve | approve | needs_approval |
list_papers and ask_hub, tools its node never declared. decide() denied both before dispatch. The run ended asking for approval, not fetching. Across the arm, 17 of 17 out-of-capability calls were denied, 0 dispatched; 13 wrote a trace row, and 4 hub_status denials were logged only (E217).4d · Asking the hub
The default came first: a loop where workers search SQLite's full text index (FTS5) and rewrite their queries until something matches. Then the question was whether another shape or retriever beat it, so four arms ran on the same 50 questions: a loop or a planned single pass, over full text or dense retrieval (search by meaning, through embeddings).
The planned pass over full text broke. FTS5 needs every word of a query to match, and a planner writes its queries before it has seen any hub text, so it answered "not in hub" on 38 of 40 answerable questions. Full text only works with a loop that can retry.[17][17]Wrongly "not in hub", of 40 answerable: planned full text 38 / 36, planned dense 10 / 11, loop full text 3 / 3, loop dense 0 / 0 (E233).[17]Wrongly "not in hub", of 40 answerable: planned full text 38 / 36, planned dense 10 / 11, loop full text 3 / 3, loop dense 0 / 0 (E233).
The loop over dense scored about the same as the loop over full text, and failed in a worse direction. Dense retrieval always returns neighbours, so the loop answered questions the hub did not cover from near misses: it declined 8 and 7 of 10 where full text declined all 10. It also cited less validly, ran out of tokens more often, and cost the most.[18][18]Loop over dense: citation validity 0.8865 vs 0.9335 over 66 citations (E244); ran out of tokens 10 / 14 times vs 3 / 4; $0.0088 vs $0.0073 per question.[18]Loop over dense: citation validity 0.8865 vs 0.9335 over 66 citations (E244); ran out of tokens 10 / 14 times vs 3 / 4; $0.0088 vs $0.0073 per question.
The planned pass over dense worked, because a blind query still finds neighbours by meaning. Dense is what makes a single pass possible. It was two behind on score in both repeats and gave up on about a quarter of answerable questions, for about a fifth of the loop's tokens.
| arm | correct of 50 | wrongly "not in hub" of 40 | out-of-hub declined of 10 | median tokens | USD per question | median latency |
|---|---|---|---|---|---|---|
| loop, full text | 42 / 41 | 3 / 3 | 10 / 10 | 47k / 51k | $0.0076 / $0.0071 | 22 / 21 s |
| loop, dense | 40 / 36 | 0 / 0 | 8 / 7 | 66k / 63k | $0.0088 / $0.0087 | 37 / 37 s |
| planned, dense | 40 / 39 | 10 / 11 | 10 / 10 | 9.1k / 9.3k | $0.0018 / $0.0018 | 26 / 28 s |
| planned, full text | 11 / 14 | 38 / 36 | 10 / 10 | 2.7k / 2.7k | $0.0008 / $0.0008 | 9 / 10 s |
The next round added a hybrid retriever, full text and dense fused. Planned over hybrid scored the same as planned over dense at the same cost and gave up more often, so planned over dense shipped as low, cheaper and not faster. Loop over hybrid at a doubled token ceiling had the highest count in the table.[19][19]low missed its pre-set limit of 4 wrongly declined questions in both repeats and shipped as an override, with a "retry with medium" hint (E272, E275).[19]low missed its pre-set limit of 4 wrongly declined questions in both repeats and shipped as an override, with a "retry with medium" hint (E272, E275).
One more run took the hybrid out of that arm: plain full text, same raised ceiling. It came within noise of hybrid, 36 and 36 of 40 against 37 and 39, at two thirds of the tokens and under half the latency, and it declined more of the questions it should. The ceiling had done the work, not the retriever. At the old ceiling, every consult that ran out of tokens had been wrong. Loop over full text at 240k tokens shipped as medium.[20][20]Loop over hybrid vs loop over full text, both at 240k: +0.05, not separated (E278); on the tuning set it declined 5 of 10 out-of-hub questions against 7 (E274). The ceiling raise alone: +0.0875, ahead (E278); at 120k, 12 of 130 consults hit the ceiling and none was correct (E277).[20]Loop over hybrid vs loop over full text, both at 240k: +0.05, not separated (E278); on the tuning set it declined 5 of 10 out-of-hub questions against 7 (E274). The ceiling raise alone: +0.0875, ahead (E278); at 120k, 12 of 130 consults hit the ceiling and none was correct (E277).
| arm | token ceiling | correct of 40 | wrongly "not in hub" of 30 | out-of-hub declined of 10 | median tokens | USD per question | median latency |
|---|---|---|---|---|---|---|---|
| loop, full text | 120k | 31 / 34 | 1 / 0 | 10 / 10 | 56k / 62k | $0.0079 / $0.0081 | 25 / 24 s |
loop, full text (medium) | 240k | 36 / 36 | 1 / 1 | 10 / 10 | 55k / 51k | $0.0086 / $0.0089 | 25 / 24 s |
| loop, hybrid | 240k | 37 / 39 | 0 / 0 | 9 / 10 | 80k / 78k | $0.0123 / $0.0118 | 52 / 60 s |
planned, dense (low) | 120k | 34 / 33 | 5 / 7 | 10 / 10 | 10.4k / 10.2k | $0.0021 / $0.0021 | 34 / 51 s |
| planned, hybrid | 120k | 31 / 33 | 9 / 7 | 10 / 10 | 10.3k / 11.1k | $0.0020 / $0.0021 | 34 / 33 s |
4e · When the hub does not know
A gap is an answer. When the judge decides the hub does not cover a question, christina says so, and then tries to say what would: real papers, with arXiv ids, ready to add.
The first version asked the model for those papers from memory. In an early run one proposal was an arXiv id for a different paper than the one it named. At scale that was the normal case: 63% of memory's proposals were wrong but fetchable, real ids for the wrong paper, which an approval would have fetched without complaint.
So the model stopped naming papers. It now writes up to three search queries; code sends them to OpenAlex, with arXiv as the failover, drops papers already in the hub, and issues an id for each result; the model then picks from that list by index. An id it did not receive from the search cannot pass decide(). With the model choosing only from search results, wrong but fetchable proposals went to 0 of 152.[21][21]Wrong but fetchable: memory 41 / 65 and 35 / 56 (76 / 121, 63%); lookup 0 / 78 and 0 / 74 (E287). Memory re-proposed papers already in the hub 4 times per repeat, lookup 0 (E289).[21]Wrong but fetchable: memory 41 / 65 and 35 / 56 (76 / 121, 63%); lookup 0 / 78 and 0 / 74 (E287). Memory re-proposed papers already in the hub 4 times per repeat, lookup 0 (E289).
Choosing the engine took a probe. The rule written before it picked arXiv, because OpenAlex missed 4 of 14 ids. The misses came from the lookup path: OpenAlex merges a preprint into its published version, so a lookup by arXiv DOI misses exactly the papers that got published. Searched by title, it found all 14. OpenAlex became the engine, against the rule, with the reason recorded.[22][22]OpenAlex by arXiv DOI 10 / 14, by title search 14 / 14; arXiv API 14 / 14 by id but 2 / 5 on search (E281, E282). arXiv takes over on a transport or quota failure.[22]OpenAlex by arXiv DOI 10 / 14, by title search 14 / 14; arXiv API 14 / 14 by id but 2 / 5 on search (E281, E282). arXiv takes over on a transport or quota failure.
Lookup shipped as the only path even though it finds little: on questions about 2026 papers it found the target 0 times in 14, and the misses happened at search, before the model could pick. On well known papers memory found more, 7 of 10 against 4 and 5. The decision rested on safety alone: a gap with no useful proposal costs me one manual christina ingest; a wrong id costs a fetch of a paper nobody asked for.[23][23]2026 papers: 0 / 14 in every arm, the target in the candidate pool 0 or 1 of 14; well known papers: lookup 4 and 5 vs memory 7 of 10 (E288).[23]2026 papers: 0 / 14 in every arm, the target in the candidate pool 0 or 1 of 14; well known papers: lookup 4 and 5 vs memory 7 of 10 (E288).[24][24]The lookup chain costs about $0.0008 and 17 s per gap (E291).[24]The lookup chain costs about $0.0008 and 17 s per gap (E291).
Then it stops. A gap with proposals suspends the run and returns one command, christina approve <run_id>, built from the server's run id and nothing the model wrote. Only I run it: it refuses without a terminal, has no --yes, and no MCP tool can resume a consult. Claude Code can relay the command; it cannot run it.[25][25]First approval round trip: 1 ingest spawned for the 1 of 3 proposals approved, fetched byte for byte; a second approve on the same run found nothing pending (E140).[25]First approval round trip: 1 ingest spawned for the 1 of 3 proposals approved, fetched byte for byte; a second approve on the same run found nothing pending (E140).
approve needs a terminal, has no MCP tool and no --yes. A person approves one paper; only then is anything fetched.4f · Cost
Cost is an output of every run, not a spreadsheet afterwards. Every model call and every hub read is a row in the trace, with its tokens, its price and the rate card that priced it, and a run's total is the sum of its rows. The table below is the answer from the clip at the top of the page, read straight out of the hub.
| node | calls | steps | prompt tokens | completion tokens | USD actual |
|---|---|---|---|---|---|
| plan | cheap model | 1 | 297 | 317 | 0.0001 |
| worker | cheap model | 12 | 70,255 | 2,417 | 0.0098 |
| worker | hub reads | 30 | 0 | 0 | 0.0000 |
| judge | cheap model | 1 | 25,103 | 614 | 0.0034 |
| synthesize | deep model | 1 | 25,144 | 2,329 | 0.0037 |
| score | hub reads | 8 | 0 | 0 | 0.0000 |
| total | 53 | 120,799 | 5,677 | 0.0170 |
Most of that answer is the worker loop: 57% of its tokens went to the cheap model calls that search, read and decide what to look at next. The deep model runs once, to write the answer. The 38 hub reads cost nothing and are still rows, each one checked by decide(); the 8 under score are the citation check behind the two warnings in the clip.[26][26]Worker cheap-model calls: 72,672 of 126,476 tokens (57%); deep model: 1 call, $0.0037 of $0.0170. Demo run 7df9d9c3, 2026-10-04, rate card 2026-09-30.[26]Worker cheap-model calls: 72,672 of 126,476 tokens (57%); deep model: 1 call, $0.0037 of $0.0170. Demo run 7df9d9c3, 2026-10-04, rate card 2026-09-30.
Tokens are the unit for comparing runs, and dollars are derived. Early on a key was on a free tier and every dollar read zero; token counts still compared. Two totals are kept apart: what the key was billed, and what the same calls cost at the rate card. A call with no price row marks the second total unknown instead of counting as zero. That is why 4a's first comparison, where Claude Code ran on a subscription, is quoted as actual against list price and never as a ratio, while the second, with both arms billed, is.[27][27]Free tier: $0 actual, $0.1005 at the rate card for the first phase's corpus (E72).[27]Free tier: $0 actual, $0.1005 at the rate card for the first phase's corpus (E72).
Since no ceiling has a default, a run cannot start without a token, step, wall clock and dollar limit, and a call that would cross one is denied before the request is sent, not refunded after it.[28][28]5 of 5 budget denials sent no prompt, and the run ended partial with finished work kept (E41).[28]5 of 5 budget denials sent no prompt, and the run ended partial with finished work kept (E41).
Across the evaluated hub, an ingest cost $0.034 and 144k tokens per paper on average, a medium question about $0.009 and a low one $0.002. Everything measured in the project, every grid and every rerun, cost at least $18.26, not counting the Claude Code sessions that built it. Every model route runs under zero data retention: the hosts keep none of the text they are sent.[29][29]Per paper: mean $0.0343 and 144,488 tokens, range $0.0214-$0.0956, n = 20 (E261). Per question: medium $0.0086 / $0.0089 (E277), low $0.0021 (E272).[29]Per paper: mean $0.0343 and 144,488 tokens, range $0.0214-$0.0956, n = 20 (E261). Per question: medium $0.0086 / $0.0089 (E277), low $0.0021 (E272).[30][30]$2.1397 for the first seven phases plus $16.12 for the rest, ≥ $18.26 actual, plus $2.758 at list price for the subscription arm (E162, E300).[30]$2.1397 for the first seven phases plus $16.12 for the rest, ≥ $18.26 actual, plus $2.758 at list price for the subscription arm (E162, E300).
5 · Limitations
The comparison in 4a is the claim this page leans on hardest, and it is the narrowest. Thirty questions per run, one hub of twenty papers, questions drafted by Claude and accepted by me. Not separated means no difference at that size; a larger or harder set could pull the arms apart. It was also run before medium's ceiling was raised, and the mode that ships has not been put against Claude Code since.[31][31]0.850 vs 0.867, n = 30 (E173); the comparison ran 2026-09-27, the ceiling was raised 2026-10-02 (E278).[31]0.850 vs 0.867, n = 30 (E173); the comparison ran 2026-09-27, the ceiling was raised 2026-10-02 (E278).
Where the arms did differ, Sonnet was ahead: more of its citations resolved to the span they named, and it answered every question that needed more than one paper. Cross paper questions are the weakest cell in every christina table on this page, so those are the answers whose citations are worth checking.[32][32]Citation validity 95.4% vs 73.9 / 84.0%; cross paper 10 / 10 vs 6 / 7 of 10 (E174, E175).[32]Citation validity 95.4% vs 73.9 / 84.0%; cross paper 10 / 10 vs 6 / 7 of 10 (E174, E175).
Every stored claim is anchored, but not every claim is stored. Against claims I marked by hand, the shipped prompts matched 27 of 66 on one paper and 30 of 49 on another, one run each. Absence from the hub is weak evidence of absence from the paper.[33][33]27 / 66 and 30 / 49 gold claims matched, n = 1 per paper (E148).[33]27 / 66 and 30 / 49 gold claims matched, n = 1 per paper (E148).
low says when it might be wrong. It gives up on about one answerable question in five, is not faster than medium, and every "not in hub" it returns carries a hint to retry with medium. The next section turns that retry into a mode.[34][34]Wrongly "not in hub" 5 / 30 and 7 / 30; median latency 34 / 51 s vs 25 s (E272, E277).[34]Wrongly "not in hub" 5 / 30 and 7 / 30; median latency 34 / 51 s vs 25 s (E272, E277).
Lookup never hands over a wrong id, and it often hands over nothing: on questions about 2026 papers it found the target 0 times in 14. When I already know the paper, christina ingest <id> is the shorter path.
Some things were never measured. Attack success is the main one. The honeypot drew one plausibly steered call, which is not a rate. An earlier attack phase never produced numbers at all: when Claude was asked to draft the attack papers and build the harness that would run them, Anthropic's safety guardrails stopped the responses, so I wrote the attacks myself and recorded that phase's attack rows as not measured. No answer was graded by a judge model; every verdict on this page is mechanical. And HTML against PDF as the source format is undecided: PDF yields more claims in total, HTML a higher share of the claims its text contains, and the two readings disagree.[35][35]Not measured: attack success (E158, E220), judged answer quality (E142). HTML vs PDF: 19-20 vs 21-24 of 55 gold claims raw, 44.2-46.5% vs 38.2-43.6% once corrected, 3 runs each (E99).[35]Not measured: attack success (E158, E220), judged answer quality (E142). HTML vs PDF: 19-20 vs 21-24 of 55 gold claims raw, 44.2-46.5% vs 38.2-43.6% once corrected, 3 runs each (E99).
6 · What is next
The next change is a mode built from the two that ship. low is cheap and gives up visibly; medium answers more at about five times the tokens. An escalating mode runs low first and reruns as medium only when low says "not in hub". If about one question in five escalates, a question costs around $0.004, a figure reasoned from the two measured modes and not measured itself. It ships after one run with its rules written first: not behind medium on the primary score, no more false "not in hub" than medium, and its latency stated, because an escalated question pays for both.[36][36]About 80% of questions at $0.002, plus about 20% at $0.002 + $0.009. Reasoned, not measured.[36]About 80% of questions at $0.002, plus about 20% at $0.002 + $0.009. Reasoned, not measured.
The second is a read tool for entities: every paper that mentions a model, a dataset or a method, with the spans. It failed its own bar before it was built. Only 45 of 1,014 entities appear in more than one paper, 4.4% against a 10% threshold set in advance, partly because "LLM" and "large language model" are still two entities. Canonical names come first, then the share is measured again.[37][37]45 / 1,014 = 4.4% cross paper entities, threshold 10%; 1.8% at 6 papers (E294).[37]45 / 1,014 = 4.4% cross paper entities, threshold 10%; 1.8% at 6 papers (E294).
Third, the trace gets a viewer. Every run already writes one row per call; exporting each run as a trace, one entry per call, to a self hosted collector would show per node waterfalls without a hand written query. This was deferred during the build and is now on the roadmap, on three conditions: local only, off by default, and never on the decision path, so a collector outage cannot fail a run and paper text never leaves through telemetry.
Two more wait on data rather than code. A high mode needs a harder question set first: medium already scores 36 of 40, which leaves no room to show a gain.[38][38]A high mode could win by at most 4 of 40, against a margin of 4 set in advance.[38]A high mode could win by at most 4 of 40, against a margin of 4 set in advance. Dense or hybrid retrieval inside the loop gets another look if the hub grows to hundreds of papers, the trigger recorded for it.
El Psy Kongroo.