Skip to content
engineeringaccuracycontinuitybehind-the-scenes

Why we're turning accounts on slowly

PlotLens was missing real continuity errors and reporting 'clean' anyway. What we found when we went looking, and why each fix was smaller than the obvious one.

Jeremy Corbello··Updated

You can create a PlotLens account today. We’ve been slow about turning them on, and that is deliberate rather than a backlog. It’s a fair thing to be impatient about, so here’s the actual reason.

PlotLens was missing things.

Not crashing. Not erroring. Missing real inconsistencies in real manuscripts, and then reporting that everything looked fine. For a continuity checker that’s the worst available failure, because it’s the one you can’t see. A tool that goes down, you notice. A tool that says “no issues found” when there are eleven — you thank it and keep writing.

We found several of these.

Three of them turned out to be sitting in a row, each one starving the next:

Where the continuity pipeline was losing information Five stages run in order: read the manuscript, pull out details, confirm each detail against the text, build the Story Bible, then retrieve relevant facts to check new pages against. Three stages were silently losing information. Confirming details discarded every detail whose claimed position was off by 65 or more characters. Building the Story Bible therefore produced zero facts. Retrieval returned none of the relevant facts under a timeline-heavy load. The final check still reported a clean result. 1 · Read the manuscript 2 · Pull out details 3 · Confirm each one in the text 4 · Build the Story Bible 5 · Retrieve rules, check new pages Dropped every detail whose position was off by 65+ characters Nothing left to build from: zero rules Returned none of the rules that mattered Result shown to you: no issues found
Three stages losing information quietly. The fifth still reported a clean result.

Here’s what each one was. (If you want the overview of what we’re building first, that’s the continuity pillar guide.)

The clean report that meant nothing

Start with the worst one, from this week.

A manuscript went through the pipeline in production. Processing completed. No errors anywhere. And the system built zero facts from it — meaning it had learned nothing about the book. Any check you ran afterward would have compared your new pages against an empty Story Bible and found, inevitably, nothing wrong.

Your Story Bible is the set of facts your story has established — what readers call canon.

A perfect score from a checker that hadn’t read anything.

It got worse when we went looking, because our own monitoring had no record of the run at all. Production drops routine log messages, and “canon inference finished” was classified as routine. A run that produces zero rules also rejects zero rules, so the rejection counter stayed silent too. Every signal we had was quiet, and quiet reads exactly like “this step never ran.” It had run. It had run perfectly and produced nothing.

The cause is worth explaining, because it’s a good illustration of how these things break.

When a model finds a detail in your manuscript — Maren’s eyes are grey — we ask it where it found that, as a character position in the text. Then we go to that position and confirm the sentence is really there. That confirmation is the whole product. It’s the difference between a citation and a guess.

A model cannot count characters. Ask one for the exact offset of a sentence inside a twelve-thousand-character passage and you get a number that is confident and approximately useless. We knew that, so the code searched a window around whatever position it claimed.

The window was 64 characters wide.

We measured the failure by taking spans we knew were correct and pushing them off by known amounts:

How far off the claimed position was Spans we could still confirm
0, 10, 50 characters 8 of 8
65, 200, 1,000 characters 0 of 8

A cliff, not a slope. One character past the window and we lost everything.

Because a detail we can’t locate is a detail we refuse to keep, the pipeline threw all of them away, built no rules, and reported success. The refusing part was correct. Refusing everything and calling it a good run was not.

Why we didn’t just widen the window

The obvious repair takes ten seconds. Make the window 500 characters. Or 5,000. Suddenly you’re finding spans again.

You’re also finding the wrong ones. A wider search doesn’t locate the quote more reliably, it locates something more reliably — and past a certain width that something is whatever text happens to sit nearby and read similarly. You’d get a citation. It would point at a real sentence in your real book. It would be the wrong sentence, and nothing on screen would tell you.

That trades a miss for a lie, and the lie is worse. A miss costs you one continuity error we should have caught. A lie costs you the ability to trust any of the ones we did catch — including the ones that depend on tracking what a character knows versus what’s true, where the difference between a real citation and a plausible one is the entire feature.

What we shipped instead: when the window search fails, search the whole document for that exact quote. If it appears exactly once, that’s where it is — a fact about your manuscript, not an inference. If it appears twice, we genuinely don’t know which one was meant, so we keep nothing.

Same test after the change: 6 of 8 recovered, at any distance. The two we still refuse are text that really does repeat.

We took a fix that recovers 75% over one that would have “recovered” 100% with some unknown share of them silently wrong. Unknown is the operative word — an error you can’t measure is an error you can’t warn anyone about.

Missing the rule that would have caught it

The second kind of miss is subtler, and it’s the one that convinced me we weren’t ready.

When you write a new chapter, PlotLens doesn’t compare it against every rule it knows; that would be enormous and slow. It retrieves the rules most likely to be relevant and checks against those. Standard approach. It also means a rule that doesn’t get retrieved cannot catch anything. The check runs, finds nothing, reports clean.

Each retrieval channel was capping its results before filtering for where you are in the story. Picture asking an assistant for the twenty pages most relevant to chapter nine, and having them hand you twenty pages chosen at random from the whole book and only then check which ones were about chapter nine. What you get back isn’t twenty wrong pages — it’s two right ones and eighteen wasted slots. Nothing further down the line can recover a page that was never handed over.

Under a timeline-heavy test load, one channel lost 16 of 16 relevant rules. Another lost 9 of 9. Another lost 25 of 25.

That’s not a percentage. Those channels were returning none of the rules that mattered, and the check sitting on top of them was reporting clean. Moving the filter ahead of the cap took those losses to zero on every channel, for the timeline-driven cases. It did not close everything: under a high-volume load the same channels can still push a relevant rule out, because they fall back to ranking by confidence and recency. That one is named, tracked, and still open — I’d rather write it down here than let “took the losses to zero” stand as though it were the whole story.

Making it impossible to say “clean” quietly

Both failures share a shape: the system did less than it claimed and told you it was fine. So we spent a stretch making sure that particular thing can’t happen without you being told.

We built a census. It takes every way the pipeline can quietly do less — a data source erroring, a source timing out, a lookup coming back empty, a result list hitting its cap — and crosses it against every place you’d read a verdict: the API, the live progress feed, the saved report, and the interface itself. Sixteen combinations. Each gets a deliberately broken run injected at that exact seam, plus a healthy control run alongside, so the test can’t be satisfied by a system that just refuses everything.

All sixteen now tell you the check was degraded instead of showing a clean result. Zero false all-clears, zero false alarms on the controls.

We also flipped a default. An internal signal the system didn’t recognize used to be treated as routine until someone classified it. Now an unrecognized signal degrades the result automatically. That produces more “partial check” verdicts than before, and it’s the right trade. A warning makes you reread a chapter you didn’t need to. An unearned all-clear makes you publish.

The boring one

Not all of it is interesting. Deleting a project’s Story Bible was taking five minutes, which made every test cycle slower and every reprocess worse. We profiled it instead of guessing. The delete itself took 22 milliseconds; a single missing database index accounted for 303 of the 305 seconds. Two indexes in the end — the guard we added to catch that class of bug promptly found a second one on its own pull request.

Most of what stands between here and opening this up properly is that, not philosophy.

How we found them

None of this came from noticing something felt off.

At the end of August we ran two independent evaluators over the whole system, each producing an architecture pass and a retrieval pass — four documents from two genuinely separate sources. Then a fifth pass whose only job was to adjudicate where those four disagreed, grade each finding by how much independent support it actually had, and throw out the ones that were a single evaluator’s hunch.

One caveat that belongs here rather than in a footnote: the consolidating pass did not re-run the evaluators against the code itself, so the code-level findings are the evaluators’ own. The report says so about itself. It would be a poor look to quote a document about not overstating things while quietly dropping the part where it limits its own confidence.

The finding that came back with the strongest possible support, identified independently by all four, was this:

A retrieval channel, extraction stage, index, or realtime path can be incomplete while the response still appears clean or valid.

That’s the thing this whole post is about. We didn’t discover it by hitting it in the wild. We went looking for exactly this class of problem, in a system that was at that moment telling us everything was fine.

The adjudicated report came out as eight numbered work packages, each with an exit condition written before the work started — not “improve retrieval” but known relevant rules survive the candidate bounds in alias, temporal, negation, and world-rule fixtures. It’s much harder to declare victory against a sentence like that.

Work package Status
1 Never report clean when the check was incomplete Merged
2 Retrieval correctness and per-channel health Merged, with the volume-load gap above still open
3 A graded corpus to measure how much we actually catch In progress — this is the gate
4 Stop re-sending the whole manuscript to the model per batch Merged
5 Make reprocessing safe to interrupt and repeat Merged
6 Make a search-index rebuild reversible Merged
7 Trace and cost a single check end to end Merged
8 Tenant isolation and adversarial-input invariants Merged

“Merged” is doing exact work in that table. It means the code landed. It does not mean I have independently confirmed each exit condition against real data — that confirmation is largely what work package 3 is for. Having just spent a paragraph praising exit conditions for making victory hard to declare, it would be a bad joke to declare it in a one-word column.

Just as useful was the list of things the report told us not to do yet: don’t change the embedding model, don’t rewrite how the manuscript is chunked, don’t add a reranking layer, don’t build a summarization tier. Not because those are bad ideas — several are probably good ones — but because every one of them is a change you can only evaluate against a recall number, and we don’t have one yet. Making them now would mean tuning a system whose omissions we don’t understand, and being unable to tell afterwards whether we’d helped.

The discipline the report actually imposed is: fix what is measurably broken, and refuse to touch what you can’t yet measure. Work package 3 is the one that unlocks the rest, and it’s the one still open.

What’s still not right

The uncomfortable part, and the reason I’d rather write this than announce a date.

When PlotLens derives a rule like Maren’s eyes are grey, it stores the excerpt the model says it came from. For the entities and relationships underneath that rule, we now check that excerpt against your document. For the rule itself, we don’t check it yet. It’s written down in the source file where it lives, in a comment that begins with the words HONEST LIMIT, so nobody here gets to forget it’s there.

And one measurement that reset our sense of how far along we were. Before we fixed the entity and relationship pointers, we counted how many of them in production actually resolved to the text they claimed to quote — 27%, not because 73% of the quotes were invented, but because the pointers were landing in the wrong place. The quoted text was real. The coordinates were off, an offset bug where a span that verified correctly on the way in failed to verify on the way back out. Nothing false was shown to anyone; we just couldn’t prove any of it, and a citation we can’t confirm is a citation we haven’t earned.

Two honest limits on that number. It covers entity and relationship citations only — the rule-level excerpt I just told you we don’t verify has never been measured at all, so I can’t give you a figure for it, only the fact that we haven’t looked. And the fix applies going forward: rows written before it were never repaired.

What turning everyone on requires

Not a feature. A number we can stand behind.

If you asked me today “what fraction of the continuity errors in my manuscript will PlotLens find,” the honest answer is that we’re still building the instrument that measures it — a graded set of real manuscript passages with known errors in them, so we can say this many instead of most of them. Every fix above moved that number. None of us knows yet what it is.

Turning everyone on before we can answer that question isn’t a growth strategy. It’s spending trust we haven’t earned, on writers handing us a book they’ve spent three years on.

When we do open it up, it’ll be because the checker catches things and we can tell you how often.

Want the detail on what it’s checking and why continuity breaks in the first place? The continuity guide is the long version, and series writing covers what changes at multi-book scale.