Skip to content
Gross Retention
Methodology

How an article gets published here

Gross Retention is written and published by machine. No person reads an article before it goes live, which is precisely why the process below exists — and why it is published rather than described in a line of small print.

The readers are institutional investors. A wrong figure here is not a typo; it is a number someone may act on. Most of what follows is machinery for deleting claims rather than for producing them.

The pipeline wakes every 30 minutes and runs the sequence below from the top. Each cycle writes at most one article and most cycles write none; the ceiling is 8 a day. Both numbers are read from the running service when this page is built, not typed into it.

Dot — what the step is
DeterministicModel judgementIndependent review
Line — what happens to the article
Moves forwardLoops backStops hereAdvisory onlyRe-enters later
SourceTopic discoveryQwen3-Next-80BDeduplicationgpt-oss-120bResearchClaude Sonnet 5 + gpt-oss-120bPremise checkGemini 3.1 ProBuildDraftClaude Sonnet 5Integration checkGPT-6 AstraEntailmentGemini 3.1 ProHeadlineClaude Sonnet 5 + Qwen3-Next-80BCheckGatesno model — deterministicSource auditGPT-6 AstraFinal readGemini 3.1 ProCold readGPT-6 AstraAdversarial readGemini 3.1 Pro + GPT-6 AstraShipPublishno model — deterministicReframethe story the sources supportRepairdelete only, never rewriteRevise → re-readsurgical edits, 2 rounds maxAbandoned — already coveredAbandoned — premise unsupportedAdvisory onlyBlockeddraft kept, with the reasonLiveverified by fetching the pageRe-enters at the gates when a fix lands
14 stages, drawn as the straight line an article takes. Repair, revision and the second opinion are the loops rather than boxes, which is why the stage list below runs longer. Four of them can stop a finished piece, two can abandon the topic before a word is written, and four arrows send work backwards rather than forwards. The straight line down the middle is the path an article takes only when nothing goes wrong — and publishing nothing is an ordinary result: a pipeline that always produces an article is one whose checks do not bind. A topic the research will not support is abandoned before drafting, which is cheaper than the alternative; a written piece that fails a check is blocked and kept with the reason attached, so a fix can be tested against the exact article that failed.

What gets refused

The rejection rate is the only number that says whether any of the machinery below binds. A pipeline that publishes most of what it proposes is not selecting, it is typing. So this is the figure the publication is managed against, and it is meant to be high: the target is to refuse more, not to hit a quota of articles. Counted from 141 topics and 21 drafts between 2026-09-04 and 2026-09-05, from the harness’s own run records.

85%

120 of 141 proposed stories were dropped before anyone wrote a word.

62%

13 of 21 pieces were abandoned or stopped at a gate after being written.

Topics

Proposed141
Already covered93
Stale — the news had aged out12
Scored below the bar13
Picked to write21

Drafts

Started21
Abandoned during research9
Stopped at a gate, never rescued4
Published8

4 of the 8 published pieces failed a gate first and were repaired by deletion. A draft is counted once here, however many times it went round that loop.

  • sources.primary3
  • style.headline_wellformed3
  • style.data_table2
  • headline.manual_review2
  • citations.minimum1
  • sources.domain_diversity1

57 articles have been published here and later withdrawn. That is a count, not a rate, and it is left as one deliberately: this ledger is newer than the archive, so there is no denominator behind it that means what a percentage would appear to mean. Publishing one anyway would be precisely the kind of unsupported figure the rest of this page exists to catch.

Stage by stage

What each box in the diagram above actually does.

Source

Nothing is written until the sources exist. A topic that cannot be corroborated is dropped here rather than written around.

01

Topic discovery

The ScoutAlibabaModel judgement

Candidate stories are pulled from live search, scored for whether an institutional investor would act on them, and ranked. Several are proposed; at most one is written.

02

Deduplication

The ArchivistOpenAIModel judgement

Each candidate is compared against everything already published. A story that restates an existing piece is dropped, even when the news is real.

03

Research

The ResearcherAnthropic + OpenAIModel judgement

Broad search, then source filtering, then full reading of what survives. Findings are cross-checked against each other before a word is drafted. A typical piece reads twenty-five or more sources and keeps around ten findings.

04

Premise check

The SkepticGoogleModel judgementCan abandon

The angle is tested against what the research actually found. If the sources do not support the story, it is reframed to the story they do support, or abandoned.

Build

The draft is written once, from the research. Every pass after it can cut, and none of them can invent.

05

Draft

The WriterAnthropicModel judgement

One piece: an opening, then sections the writer titles itself. The writer works from the research and never sees a search engine, so nothing can enter the article that was not fetched and read first. Technical language is allowed, and anything a literate non-specialist would not know carries a short explainer.

06

Integration check

The Fact-checkerOpenAIModel judgement

Sentences the research contradicts are cut. They are not rewritten — a model asked to repair a sentence after the research is finished will invent one that reads exactly like the rest.

07

Entailment

The AuditorGoogleModel judgement

Every claim is tested against the research that produced it. Claims the sources do not carry are deleted, and a figure whose citation was removed goes with it rather than being re-sourced from memory.

08

Headline

The Sub-editorAnthropic + AlibabaModel judgement

Written last, from the finished article rather than from the original commission, then checked against both the research and the body. A headline has to name its subject well enough for a stranger to place it.

Check

Two different questions, asked separately. The gates check shape — is there a citation in this paragraph, a document at the end of this link, do the tags balance — and cannot tell you whether the number in the sentence is the number on the page. The audit checks facts: every cited page is fetched and the sentence graded against it. “It passed the gates” is never the same sentence as “it is right”.

09

Gates

The RulebookDeterministicCan block

Around thirty deterministic checks, no model involved. Every citation resolves to a document. Every paragraph with a material figure carries a source. Links are intact, tags balance, no figure is stated twice, no percentage floats without its base.

10

Repair

The MechanicDeterministic

A failed gate gets one narrow, mechanical repair. A repair may copy or delete. It may never compose.

11

Source audit

The VerifierOpenAIModel judgementCan block

Every cited page is fetched and read, and each sentence carrying both a figure and a citation is graded against the page it points at. A sentence its own source contradicts is cut.

12

Comparability check

The StatisticianGoogleModel judgementCan block

Every comparison in the piece is checked against the page it cites: were these two figures measured the same way? Period, methodology, scope, basis, or a setting named nowhere in the sentence — any of them can differ while both numbers stay perfectly true. A comparison has a hidden third term, and this is the only check that looks at it.

13

Recency check

The ClockOpenAIModel judgementCan block

The figures the article leans on are searched for again, to ask whether a newer value has been published since. It reports and never substitutes: a search snippet is not a source. A stale figure is the worst kind of wrong claim to catch by eye, because it has a real source, an accurate quotation and a working link.

14

Final read

The Standards editorGoogleModel judgementCan block

The headline and the standfirst are checked against the finished body — the two sentences most likely to promise more than the article delivers.

15

Cold read

The StrangerOpenAIModel judgement

The opening is read by something told to arrive knowing nothing: is the subject placeable, is the jargon legible, is the reason it matters findable. Advisory — it reports and does not stop the piece.

16

Adversarial read

The AdversaryGoogle + OpenAIIndependent reviewCan block

Two models, from two labs that are neither of them the writer's, read the finished article as an argument and look for the failure the other checks structurally cannot see: a piece that contradicts itself, or asserts something load-bearing that nothing supports. They do not find the same things, which is the argument for running both.

17

Blocking-finding confirmation

The Second OpinionOpenAI + GoogleIndependent review

Before a finding is allowed to stop a publication, it is put to a reviewer from the other lab, with the article, and has to be agreed. The model that raised a finding is the one least likely to overturn it. This step fails closed: if the second opinion cannot be obtained, the finding does not block.

18

Revision

The Copy editorAnthropicModel judgementCan block

The one place a model is allowed to change a finished sentence, and it is fenced: it may delete the sentence, weaken it to what the sources carry, or mark it as the publication's own inference. An edit is applied only where its quoted text matches the article exactly, and any edit that would introduce a figure the original did not contain is refused and logged. Then the article goes back for another full adversarial read — at most two rounds, after which anything still serious holds the piece.

Ship

19

Publish

The PressDeterministic

The locked deploy described under The plumbing below: build, restart, then fetch the live page and confirm what a reader actually gets. Publication is recorded only after that fetch succeeds — a deploy that fails is not a publication.

The words that decide things

Three definitions carry most of the weight above. They decide which sentences get checked and which findings can stop a piece, so they are worth stating exactly rather than leaving to be inferred from the checks that use them.

A currency amount, a percentage, a number written out with percent, bps, million, billion or trillion, or a comma-grouped thousand. A bare small integer is not one, which is what keeps “Q2” and “GPT-6” from dragging a paragraph into the citation requirement. This is the definition the gates use to decide which paragraphs need a source, and the audit uses to decide which sentences to grade.

It is also the single line here most worth testing in both directions. This pattern once required a word character after the percent sign, so 30%x was a figure and 30% was not — every bare percentage in the archive was invisible to both checks, on a publication whose numbers are mostly share, growth and margin. Nothing failed. The checks simply had nothing to look at.

The reviewer that raises a finding grades it, against one test: publishing this as written would mislead a reader about money. Below that, serious is a real defect in the argument that should be resolved before the piece runs, and minor is worth an edit but would not by itself hold the piece. Only blocking and serious findings stop anything.

A blocking finding does not act on its own: it goes to a reviewer from the other lab with the article attached and has to be agreed. That step fails closed in the direction of publishing — if the second opinion cannot be obtained, the finding does not block — because the alternative is a pipeline that stops on any outage, and because the finding it was built for turned out to be one the article never made.

Every claim is checked against a page fetched during the run, and the audit refetches those pages nightly for as long as the article is live. That means a later disagreement has two possible causes — the article was wrong, or the source changed — and they are not treated alike: a figure the current page contradicts is corrected against the current page, and a premise that no longer holds is a retraction rather than a swap. Wrong figure, correction. Wrong premise, retraction.

Who is doing the work

Every stage above is a separate call to a model, given one job and no memory of the others. Naming them is not decoration: it is the only way a reader can see how much of this pipeline is one company’s judgement. Regenerated from the running configuration on every deploy, so this list cannot quietly go stale.

anthropic/claude-sonnet-5

Reporting, drafting, the headline, and the surgical fix an adversarial finding asks for.

gemini-3.1-pro-preview
gpt-6-astra

Every pass that inspects what the writer produced. A different lab, on purpose.

gemini-3.1-pro-preview
gpt-6-astra

Two independent readers, and the cross-lab second opinion on anything that would stop publication.

openai/gpt-oss-120b-Turbo
Qwen/Qwen3-Next-80B-A3B-Instruct

Jobs with a right answer: is this the same story, does this label match its link, does this term need a gloss. A frontier model here is overpaying, and open weights mean this tier can be run by anyone without a frontier account.

One rule decides the table below: whoever writes does not grade. Until 5 September 2026 the draft, the entailment check, the source audit, the final read and the cold read were the same model. The thing that wrote the article was marking its own work at four separate steps, which is not a check at all — it is the same model agreeing with itself and calling the agreement verification.

Now the writer is one lab and every check on what it produced is another. The two adversarial readers are a third and a fourth reading, from labs that are neither the writer’s. When one raises a finding serious enough to stop publication, a reader from the other lab has to agree before it can.

This is deliberately not a claim that any one model is better. It is a claim that a reviewer with a different training run has different blind spots, and that is the only property being bought here.

The handoff, in order

No lab holds two stages in a row. Six checks sharing one model would share its blind spots — six reads that look independent and behave like one. Alternating means anything the previous stage could not see meets a different lab immediately. The lab is named on every card in the stage list above: read them downwards and no two touching stages carry the same name. That is the whole claim, and it is checkable without taking our word for it.

How the searching actually works

Search is not a model, and the distinction matters more here than it sounds. A model asked what it knows about a company answers from training data that is months stale and will not say so. Search returns URLs and writes nothing; the pages behind those URLs are fetched and read in full, and a source that cannot be fetched is dropped rather than summarised from its snippet. Everything this publication asserts comes from a page fetched during the run.

Google, reached through the Gemini CLI, is the primary search. It runs the open-ended questions — what has been reported about this company, this quarter, this filing — and it is budgeted: there is a hard daily query ceiling, and each cycle reports how much of it is left. Search is the cheapest part of this pipeline to overspend on.

Serper is the failover, and it owns one thing outright: site: queries. When the question is “what does this company’s own investor-relations site say”, or “is this figure in the actual SEC filing”, that is a Serper query. It is also what runs when the primary search returns nothing usable, so a quiet failure upstream does not silently become an article with no sourcing.

After publication

Publishing is not the end of the loop. Every live article is re-audited nightly against its own sources, indefinitely — a citation that was accurate in September can point at a page since edited or withdrawn.

Live articleNightly auditevery cited page, re-fetchedStill standsno changeCorrectionfigure for figure, recordedDeletion, then retractioncut the sentence; 410 if it cannot standcorrected article goes back into the live set, and is audited again
All of it runs on a timer, including withdrawal. The ladder stops at the first step that works — substitute the figure, else cut the sentence, else pull the piece — and no step composes a word. A findings file nobody acts on is decoration, so the job acts on its own findings; the safety is in what it is allowed to do, not in whether it asks.

Corrected

A published figure that does not match its source is replaced with the one that does, figure for figure, and the correction is recorded on the article. Anything more than a swap is refused, because anything more is composition. The corrected article goes back into the live set and is audited again.

Retracted

A piece whose premise does not hold is withdrawn, and its URL returns 410 Gone rather than a soft 404, so it leaves the index instead of lingering as a dead link.

The plumbing

Everything above describes judgement. None of it moves an article on its own. What actually advances a piece from a topic to a live page is ordinary, deterministic software: the processes below, a lock, and a rule that nothing may fail silently. It is the half usually left out of a diagram, and the half that decides whether any of the above actually runs when nobody is watching.

The writing loop: discover, research, draft, check, publish. Restarts automatically if it dies. A cycle that produces nothing is a normal cycle.
Takes a lock, builds, restarts and then fetches the live page to confirm what a reader gets. Build and restart are never separated — in between, the running site serves 404s for its own stylesheets. A deploy that fails is not recorded as a publication.
Re-reads every live article against its sources, then acts on what it finds: substitute the figure, else cut the sentence, else withdraw the piece. Unattended, and it composes nothing at any step.
Fails if the publication stopped publishing. It exists because this site once printed nothing for four days while every process reported success — the harness drafted an article every cycle — six-hourly at the time — and the gate rejected each one, which is a normal exit code. Silence needed its own alarm.
Checks the publication's claims about itself. Regenerates the model manifest and fails if it disagrees with the running configuration, fetches this page and confirms it still names those models and this cadence, checks every live article and the home page serve, checks withdrawn pieces still answer 410, and runs the test suite. It exists because the three stalest claims on this page were found by two outside models reading it, not by anything here.
Serves the site on a loopback port behind nginx. Restarts on failure.
Two nginx maps decide what a dead URL answers: 410 for a withdrawn piece, 301 for one that moved. Generated from git history and a hand-kept list, never edited directly, and validated before they are installed — a bad map takes the whole site down on the next reload.

Every scheduled job here is wired to one failure handler that writes to a log, tags syslog and sends mail, rate-limited so a broken job cannot spam. This matters more than it sounds: the journal on this machine keeps about eight hours, so a job that failed overnight has left no trace by morning. A pipeline that runs unattended is only as trustworthy as its ability to say when it stopped.

The contrarian rules are different

One section of this publication argues against the prevailing reading of a story. That is the most valuable thing here when it is right and the most damaging when it is not, so it carries constraints the other sections do not.

A hard quota: one a day

At most one contrarian piece runs in a day, against a ceiling of 8 articles total. When the quota is filled, contrarian candidates are removed from the pool outright — and if nothing else is left, the cycle publishes nothing rather than filling the slot. A publication that is contrarian every day is not contrarian; it is just a different consensus.

The premise check binds hardest here

A contrarian angle is a thesis before it is a story, and the research usually supports something narrower than the thesis. When it does, the piece is rewritten to the narrower story or dropped. A commissioned angle asserting a benchmark 'is the benchmark that failed' was refused on exactly this ground and reframed to what the sources carried: a 36-point gap that came from evaluation harness design.

The disagreement has to be with evidence, not with mood

Against the dominant reading is a finding. Against the dominant feeling is a column, and this is not one. The claim has to be checkable, and the thing it contradicts has to be quotable.

It is marked on the page

A contrarian piece carries its own section mark and a different kicker colour, so a reader can see they are being offered an argument against the consensus rather than a report of it, before they have read a sentence.

The inference rule under Where this can still fail binds hardest here: a contrarian piece is mostly inference, and every sentence of it has to be marked as one.

Where this can still fail

The source audit only examines sentences that carry both a figure and a citation. A confident qualitative claim with no number and no source attached is invisible to it. That gap is the reason the adversarial read exists, and it is the gap most likely to produce an error here.

Automated checking is good at catching a claim that disagrees with its source and much weaker at catching one that no source addresses at all. So judgement calls are marked in the text rather than written as fact: the phrase “Our read is that…” appears throughout this publication and it is load-bearing. It marks the sentence as the publication’s own inference rather than something a source establishes. An inference is allowed to be an inference; what is not allowed is an inference wearing the clothes of a reported fact.

If you find an error, corrections@grossretention.com.