The image analysis pipeline
The tools agent’s registry exposes image analysis as one tool. Internally it is 12 discrete checks running in a fixed, deterministic order — no model chooses among them, and none of them is separately callable. This page documents each one in the sequence it runs.
The organising fact is that finding matching pixels is the easy half. A journal logo, a scale bar, a shared axis, a TEM instrument overlay and a reused plot template all produce perfect matches, and all of them are innocent. Nearly every stage after the detection tiers exists to refuse one measured class of false positive, and each carries the corpus evidence that put it there.
This pipeline is quarantined. It runs on every paper and reports its statistics, but its verdict is withheld — delivered severity is always indeterminate, with the uncalibrated grade preserved in a separate field. It will stay that way until the pre-registered calibration criteria are met. Nothing here should be read as an accusation against any paper.
The registered tool
One registered tool, but internally a fixed sequence of discrete checks: harvest the panels, screen out the publisher's own furniture, then four detection tiers ordered from strongest evidence to weakest, each handing the next only the pairs it did not already claim. Between them sit the gates that decide what a match is allowed to mean — and those gates are most of the engineering, because the failure mode here is not missing a duplication, it is calling a shared axis label one. The whole module is QUARANTINED: it runs, its statistics are reported, and its verdict is withheld until the calibration criteria pass.
Panel duplication (the pipeline)
analyze_paper_imagesWhen it applies
A paper with embedded raster figures — micrographs, blots, gels, photographs — and its supplements.
How it works
Runs the twelve stages below in a fixed order and merges their output into one result. It is a deterministic subprocess running alongside the statistical sweep, never called by the model: a language model mid-analysis must not be launching a computer-vision pipeline, and its findings reach the editorial stage by injection, already quarantined.
Inputs
- pdf_path
- PDF file
- si_pdfs / paper_diroptional
- supplement pathsOmitting them is recorded explicitly — a main-PDF-only run that found nothing is not “we checked the supplement and it was clean”.
- captionsoptional
- list of figure captions
- stated_reuseoptional
- true / falseThe paper says these panels re-show earlier ones; the result is then capped at indeterminate.
- furniture_sha256optional
- set of content hashesFrom the corpus furniture screen. Empty hashes are filtered out — an empty hash is the harvest's “unknown”, not a value, and testing membership naively would withhold the entire pool whenever the image library is absent.
- segment_montagesoptional
- true / falseOff by default: segmentation triples the corpus fire rate, mostly via sub-panels sharing a plot template.
Output
Duplicate groups ordered strongest tier first, each with its match geometry and the reason it survived every gate — plus panel and skip counts per document, and per-document truncation flags. Coverage is embedded rasters only: vector-drawn charts and flattened page scans are invisible, so zero groups on a vector-figure paper is coverage absence, not cleanliness. Within-paper only; nothing is compared against other papers or any external index. QUARANTINED — delivered severity is always indeterminate, with the uncalibrated grade preserved separately.
Reference
Bik, E. M., Casadevall, A., & Fang, F. C. The Prevalence of Inappropriate Image Duplication in Biomedical Research Publications. mBio 7(3), e00809–16. 2016.
Getting the panelssteps 1–3
Before anything can be compared, the pictures have to be found — and the publisher's own furniture has to be taken back out, or every paper in an imprint matches every other one.
Panel harvest
step 1 ofanalyze_paper_imagesWhen it applies
First, on the paper and every supplement that will be compared.
How it works
Pulls every embedded image stream out of the PDF, recording for each panel its source document, page, placement rectangle, content hash and a perceptual hash. A pdfplumber fallback covers documents the primary extractor cannot open; it keys panels by stream hash rather than by PDF object number. Supplement containers (.docx and .epub zips, loose image files) are harvested into the same pool.
Inputs
- pdf_path
- PDF file
- max_panelsoptional
- integerA per-document budget. Truncation is announced per document, because a truncated supplement and a truncated main PDF bound different claims.
Output
The panel pool, plus counts of what was harvested, skipped and truncated. The supplement scan reads an allowlist of subdirectories rather than walking the tree: the harvest writes its own PNGs into the paper directory, and a general recursion would re-ingest them and manufacture a byte-identical match on every paper in the corpus.
Montage segmentation
step 2 ofanalyze_paper_imagesWhen it applies
Opt-in. A six-panel composite figure is a single embedded stream, so duplication between its own sub-panels is invisible unless the composite is split.
How it works
Finds the gutters — whole rows and columns of near-white pixels — and splits on them, but only when the bands form a REGULAR grid. Regularity, not the mere presence of white lines, is what licenses the split. Columns are measured inside the row bands, because a full-width coloured banner otherwise stops any column reaching the purity threshold. A fallback path exists for montages whose border ring is coloured, which defeats the primary background estimate.
Inputs
- segment_montagesoptional
- true / falseDefault off.
Output
Sub-panel crops added to the pool, each recorded as a crop of its parent so later stages can tell a crop from an independently embedded image. Measured: turning this on raises the corpus fire rate from 5.0% to 15.1% of papers, mostly through sub-panels that share a plot template rather than content — which is why it is opt-in and why the clique demotion below exists.
Corpus furniture screen
step 3 ofanalyze_paper_imagesWhen it applies
Across a whole corpus, before the per-paper tiers run. Necessarily two-pass — the set is only knowable corpus-wide.
How it works
An image stream that appears in three or more DISTINCT papers is a journal template, society mark or advert, not a result. No within-paper rule can see this: five Bentham papers in the reference corpus each place one full-page advert twice, via two separate PDF objects — which is exactly the shape of a paper reusing an image as two results.
Inputs
- corpus panel index
- panel records across many papers
- minimum papersoptional
- integerThree. Below that, a repeat is not yet evidence of a template.
Output
The excluded hash set plus a report of exactly what was excluded and why, so the exclusion is auditable rather than a quiet absence. Panels carrying an excluded hash are withheld from every tier with a named reason.
The four detection tierssteps 4–7
Ordered strongest evidence first, each tier handing the next only the pairs it did not already claim. Tier 1 rests on the PDF's own assertion that two pictures are one object; tier 4 rests on recovered geometry. They are not interchangeable, and a finding should always be read with its tier.
Tier 1 — repeated placement of one stream
step 4 ofanalyze_paper_imagesWhen it applies
Always. The strongest tier, because the assertion of sameness is the PDF's own.
How it works
Finds streams the document itself places more than once. This is not a similarity judgement: the file says these two pictures are one object. Either the PDF object number or the stream hash supplies the identity — requiring the object number made this tier silent on exactly the documents the fallback harvest exists to cover.
Inputs
- panel pool
- harvested panels
Output
Groups of repeated placements. Placements inside captioned figure regions are evidence; reuse living entirely outside figure regions is separated out and capped at indeterminate — page decorations that survived the corpus screen (which needs three or more papers) land there rather than in a finding. Four or more placements of one stream is itself treated as a furniture signature.
Tier 2 — identical original bytes
step 5 ofanalyze_paper_imagesWhen it applies
Distinct panels whose stored bytes hash identically — the same picture embedded twice as two separate objects.
How it works
Compares content hashes of the ORIGINAL embedded bytes. The provenance gate acts here and is the point of the tier: a panel whose pixels we re-rendered ourselves never enters, however its hash falls, and the exclusion is recorded by name. Byte identity is a statement about the authors' files; a re-render is a statement about our renderer.
Inputs
- panel pool
- harvested panels
Output
Groups of byte-identical panels, and a named list of panels excluded for carrying re-rendered rather than original pixels.
Tier 3 — perceptual hash clustering
step 6 ofanalyze_paper_imagesWhen it applies
Panels that are visually the same picture but not byte-identical — recompressed, resized, or re-exported.
How it works
Computes a perceptual hash per panel and takes connected components of panels within a Hamming distance of 4. Pairs already sharing a non-empty content hash are skipped, because byte identity owns them. That emptiness test is load-bearing: a montage sub-panel carries an empty content hash by contract, so a naive equality check read every pair of crops as “the same bytes” and skipped the entire tier for them.
Inputs
- eligible panels
- panels passing the texture and glyph gate
Output
Clusters of perceptually near-identical panels. A panel needs only a perceptual hash to enter, not a saved image file — demanding a file made every montage crop inert, excluded from the perceptual tiers and structurally barred from the byte-identity ones, i.e. compared by nothing while the record claimed otherwise.
Tier 4 — keypoint match and geometric verification
step 7 ofanalyze_paper_imagesWhen it applies
Pairs the first three tiers did not claim — the tier that catches a region rotated, rescaled, mirrored or cropped out of one panel and into another.
How it works
Detects up to 2,000 keypoints per panel, matches descriptors under a 0.6 ratio test, then fits a partial-affine transform with RANSAC at a 3-pixel reprojection tolerance. A match must survive on geometry, not on match count alone: at least 30 inliers, an inlier ratio of 0.25, a recovered scale between 0.2× and 5×, shear under 15°, and keypoints spread over at least 15% of the panel rather than piled in one corner.
Inputs
- candidate pairs
- panel pairs with loadable pixels
Output
Verified matches with their recovered geometry — scale, rotation, shear, translation and the inlier bounding box — which the classifier and grain stages then interrogate. A panel whose pixels cannot be loaded loses THIS tier only; its perceptual hash still took part in tier 3, and the record says so rather than implying it was never compared.
The gatessteps 8–12
Most of the engineering. The hard problem is not finding matches — it is refusing the ones that mean nothing: a shared axis label, an instrument overlay, a plot template, a figure that moved pages during typesetting. Every refusal is recorded by name, because a panel excluded here was not examined and cleared, it was not examined.
Texture and glyph gate
step 8 ofanalyze_paper_imagesWhen it applies
Before the perceptual tiers, on every panel, deciding which are eligible at all.
How it works
Two refusals with different reasons. A panel under 128 pixels on its longest side is treated as a rendered letter or glyph — a panel label, not a result. A panel whose grey-level entropy falls below 3.5 bits, or whose top two intensity bins hold more than 90% of its pixels, has too little texture for perceptual matching to mean anything; western blots live here and wait on a separately calibrated channel. Order matters: the glyph trap is judged first, so a rendered letter is named as what it is rather than as a low-texture blot.
Inputs
- panel
- a harvested panel
Output
Eligible, or a named exclusion reason carried into the result. This is why a null result must be read alongside the skip counts: panels excluded here were not examined and cleared, they were not examined.
Match geometry classifier
step 9 ofanalyze_paper_imagesWhen it applies
On every keypoint match, before it is allowed to count as a duplication.
How it works
Answers two independent questions from the recovered geometry. Is it FURNITURE — a long thin band (a caption strip, running header, axis or scale bar), at aspect ratio 10 or more wherever it sits; or an elongated element at aspect 5 or more that has not MOVED between the two panels, which makes it part of the frame rather than content carried across. Is it PARTIAL — a region duplication covering under half of BOTH panels, in which case the grain test is disqualified. Furniture is never decided by area: the areas of the real and spurious cases overlap and these two signals do not. A 2% area floor is only a backstop for a match too small to be anything at all.
Inputs
- match geometry
- recovered transform and inlier statistics
Output
A furniture flag and a partial-region flag. Measured: the static-element rule alone reclassified all 45 groups of one paper, a TEM instrument info bar recovering a translation of exactly 0 px at aspect 7.6, where every real duplication in the corpus moved between 278 and 1,057 px. The partial exemption additionally demands 80 inliers, well clear of the 30 a match needs to exist — because the exemption disqualifies grain, and geometry must then carry the finding alone.
Grain residual test
step 10 ofanalyze_paper_imagesWhen it applies
On whole-panel matches, to separate the same pixels from the same subject photographed twice.
How it works
Aligns one panel onto the other using the recovered transform, subtracts the low frequencies, and correlates what is left — the sensor grain. Genuinely identical pixels carry identical grain; two photographs of the same specimen do not. The measurement is restricted to the bounding box of the matched keypoints rather than the whole warped overlap.
Inputs
- aligned panel pair
- two panels plus their transform
Output
A correlation coefficient, or null WITH the reason it could not be computed — a discriminator failing silently would leave a group looking merely unclassified, which reads as “we checked and could not say” when nothing was checked. Restricting to the matched region is measured, not cosmetic: on one real case the same pair scores 0.506 over the whole overlap and 0.718 over the matched region, because averaging across surrounding non-duplicated tissue dilutes the statistic. Grain answers “are these the same pixels”; asking it about pixels the match never claimed is the wrong question, which is why the partial-region exemption exists.
Sibling-clique demotion
step 11 ofanalyze_paper_imagesWhen it applies
When montage segmentation is on and sub-panels of one figure match each other.
How it works
If every pair among a figure's sub-panels matches, that is a shared plot template, not duplication. Completeness is the signal: real reuse between two condition panels of one figure is a sparse pair, not a complete graph.
Inputs
- match groups and sibling sets
- groups plus each figure's sub-panels
Output
Demoted groups, named as template matches. Measured on one paper: 66 groups, exactly the 66 pairs of 12 sub-panels. That is the template. Sparse matches among the same siblings are left standing.
What this pipeline cannot see
- Vector artwork. Coverage is embedded rasters only. A chart drawn as linework, or a page flattened to a scan, carries no image stream to harvest — so zero findings on a vector-figure paper is coverage absence, not cleanliness. Read the panel and skip counts before reading the verdict.
- Anything outside this paper. Every comparison is within one paper and its supplements. Nothing is checked against other papers, other authors, or any external image index.
- Low-texture panels. Western blots and similar flat-field images are excluded from the perceptual tiers by design and wait on a separately calibrated channel — precisely the material that most often matters.
- Legitimate re-use, unless it is declared. A representative micrograph re-shown and said to be re-shown is not evidence. The caption scan and the stated-reuse flag discount those cases, but a paper that reuses a panel without saying so is indistinguishable, at the pixel level, from one that does so deliberately.
Back to the full toolkit, or see the findings the toolkit has reported.