The Live Variant Playground
Shorkie · Chao et al. 2025 · fold f0, 14,253,567 parameters · act 9 crosses all 8Language model →LabPost1The system and the question
A model with 5,215 outputs has no single "prediction", and an interpretability page that never says which one it explains can contradict itself without any of its numbers being wrong. So this act fixes the object: one 16,384 bp window, one gene's own output bins, one track subset, one scalar. Everything below differentiates, perturbs or ranks against that number — pick a locus and a region in the bar above and read straight down.
The window, and the one number every method differentiates
- Question
- What does the model take in, and what single scalar does every method on this page differentiate?
- Why this matters
- Every panel below differentiates, perturbs or ranks against one number. If that number is not fixed and stated, two panels can disagree because they were quietly explaining different quantities — the commonest way an interpretability page contradicts itself.
- Method
- A 16,384 bp window centred on one gene. f(x) = log2( Σ over that gene's bins of the mean over the 384 T0 RNA-seq tracks, + 1 ) — the paper's own logSED target, ensemble.py:97-104.
- Intuition
- A model with 5,215 outputs has no single 'prediction'. Scalarising to one gene's own bins over one declared track subset is what makes 'which bases drove it' a well-posed question at all.
- Cost
- Free. The sequence and the gene models ship with the page.
- What would refute it
- Nothing here is a claim; it is the definition everything below is measured against.
What the answer tells us Nothing is claimed here; this is the definition. Its value is that every number below is a change in this one quantity, so where two panels disagree they are disagreeing about the model and not about what was measured.
Figure 4 motifs. Database consensus matches found in this window on either strand — not the paper's own annotations, which this scan neither reproduces exactly nor claims to. Click one to scramble it and re-run: same length, same base composition, only the order destroyed, so any change is attributable to the arrangement rather than to GC content. The effect is measured over this window's own gene, not the whole window — a 14,336 bp yeast window holds a dozen genes and the tallest is usually not the one whose promoter you edited.
Expect a wide range. Across these six windows the splicing motifs dominate — knocking out DTD1's 5′ splice site costs 34% of its predicted expression and its branch point 21% — while transcription-factor sites range from −18.3% (HOP2's TATA box) through −10.5% (KRE33's Reb1) and −7.4% (FUN12's RRPE) down to sites that move nothing at all. A near-zero result is a real answer about that site, not a broken button.
Each locus is a 16,384 bp window — the model's full context — centred on the named gene. Six of them are the windows Figure 4 of the paper prints, at the published coordinates; the rest are classic high-expressers and the galactose regulon. In goes 16,384 × 170 channels — 4 DNA, 1 unused, 165 species — and out come 896 bins of 16 bp across 5,215 tracks, with 1,024 bp cropped from each end.
Every method below differentiates the same scalar, and that is the only reason they can share an axis or be compared at all:
f(x) = log₂( Σgene bins mean384 T₀ tracks coverage + 1 )
That is the paper's own logSED target (ensemble.py:97-104): the mean is over tracks, the sum is over bins, the pseudocount sits inside the log, and the base is 2. The order matters — a mean of logarithms is not the logarithm of a mean. Being a log ratio is what makes a silent promoter and a maximal one comparable, and it is why “a stronger effect” on this page is meaningless without saying stronger relative to what. The 384 _T0_ tracks are the uninduced baseline subset the paper's Figure 4 uses, not the whole 3,053-track RNA-seq block.
2What can the model see at all?
Before asking which bases drove a prediction, ask how many bases could. The architecture reaches the whole window; "can see" and "depends on" are different claims and only the second is measurable. This scopes every method that follows.
Before asking which bases drove it — how many bases can? effective receptive field
- Question
- How much of the 16,384 bp of context does the prediction actually depend on?
- Why this matters
- Every attribution below is scoped by what the model actually uses. If a prediction is settled by the nearest kilobase, an attribution peak six kilobases away is noise — and a method that reports one is wrong rather than interesting.
- Method
- Keep a centred core of real sequence, dinucleotide-shuffle everything outside it, and grow the core until the prediction settles inside ±5%.
- Intuition
- Architectural reach and dependence are different claims: the transformer connects the whole window, but 'can see' is not 'uses'. Shuffling a flank while preserving its dinucleotide composition removes the information without removing the sequence.
- Cost
- 5 shuffles × 8 radii = 40 forward passes a locus, 23 windows.
- What would refute it
- Two controls. At the largest radius the window is entirely real, so that point must reproduce the full-context prediction — and convergence requires every larger radius to stay inside the band too.
What the answer tells us This is the practical scope of every attribution on the page. A peak outside the radius that settles a locus is not a long-range dependency — it is a base the prediction does not use, and reading it as regulatory would be reading the method.
Every attribution method on this page distributes credit across all 16,384 bases, and the architecture really can reach all of them — dilated convolutions into full self-attention at the bottleneck. But can see and depends on are different claims, and only the second is measurable. Keep a centred core of real sequence, replace everything outside it with dinucleotide-shuffled DNA, and grow the core until the prediction stops moving.
The effective context is a property of the gene, not of the model. The median window is settled by ±2,048 bp of the ±8,192 available, and the six constitutive glycolytic genes — TDH3, PGK1, ACT1, ADH1, FBA1, PDC1 — are all finished by ±1 kb: a housekeeping promoter determines its own gene. GAL1 needs the entire window, and GAL3, FUN12 and KRE33 need ±4 kb. That is the practical scope of every attribution on this page — outside a window's own radius there is little left to attribute.
Why shuffled and not zeroed, and the two controls
Zeroing the DNA channels is what occlusion does, and it is indistinguishable from a run of N — a sequence the model has never seen. The answer would then be about out-of-distribution input rather than about context. A dinucleotide shuffle is the right null because yeast promoters carry strong dinucleotide bias, poly(dA:dT) above all, and a mononucleotide shuffle destroys that too. The shuffle is Altschul-Erikson — a random Euler path through the dinucleotide graph — and make_receptive.py refuses to run unless all sixteen dinucleotide counts survive it exactly on a 4,000 bp probe.
At the largest radius the window is entirely real, so that point must reproduce the full-context prediction. It does, at every locus, which is the check that the splicing and scoring path is not itself perturbing anything. And convergence requires every larger radius to stay inside the ±5% band as well, so a curve that crosses once and wanders back out is not counted as converged.
5 shuffles per radius per locus, 8 radii, 23 windows. The spread across shuffles is stored alongside each point.
3What does it predict?
Before any explanation is worth reading, the thing being explained has to be visible and comparable. This act draws the prediction itself and every lane that explains it on one zoomable axis with everything that explains it. Zoom in far enough and the per-base attribution lanes turn into DNA logos in place; every lane is the same base pair at the same x.
Predicted coverage — the full model this window, 16,384 bp
- Question
- What does the model predict across this window, what drove it, and what is actually annotated there?
- Why this matters
- This view replaced four stacked panels that each computed their own left inset, so one base pair sat at four different x positions and no two lanes could be compared by eye. A shared axis is what makes 'these methods disagree here' a statement about the methods.
- Method
- One forward pass for the curve (896 bins × 16 bp); every attribution lane rc-averaged and drawn at its own real step; annotation from SGD, ORegAnno, Harbison/MacIsaac and a JASPAR scan.
- Intuition
- Every lane answers a different question about the same coordinates. Drawing them on one axis at each method's own real resolution lets a divergence be read directly instead of inferred across two figures at different scales.
- Cost
- Free. Every locus was run offline and the packs ship with the page.
- What would refute it
- This is a drawing, not a claim. Its correctness condition is that one base pair lands at one x in every lane — which one canvas makes true by construction rather than by assertion.
What the answer tells us Where the lanes part company is the finding, and on one axis it is readable directly rather than inferred. A method drawing a peak no other method draws is either resolving something the others cannot, or reporting its own estimator.
Everything on this page is about ONE window at a time — the locus in the bar above. This is what the model predicts across it, what drove that prediction, and what is actually annotated there, on one shared axis so a peak can be read against a gene. For the same quantities across all 12,157,105 bases, and for the language model beside them, the genome browser is the other half of this lab.
Drag to pan, drag the ruler to select, scroll to zoom — and keep going. Past about seven pixels a base the three per-base attribution lanes stop being bars and become DNA logos — mutagenesis, gradient × input and integrated gradients, each in place, each still the same lane — and the sequence appears beneath them. Nothing new pops up: the same three measurements are simply drawn where every base has room for its own letter. Occlusion and attention rollout stay as bands, because 64 bp and 128 bp a value cannot be letters and stretching them into some would claim a resolution neither has. Every lane is the same base pair at the same x, so a peak can be read straight down against a motif and a gene. Click anywhere to trace the 16 bp bin under the pointer; the shaded band is what the region-conditioned panels below are keyed on — gradient x input, integrated gradients, occlusion and rollout, but NOT mutagenesis, which keeps its own focal-gene target and is unconditional. The flanks are shaded because the head crops 1,024 bp from each end — the curve covers 896 bins of 16 bp and the model still reads the whole 16,384 bp to produce them.
Every attribution lane is signed, and that is the whole reading. A bar above the rule is a base whose presence RAISES the prediction, below it one that lowers it. Each lane is drawn at its own real step and labelled with it: mutagenesis and the two gradient methods resolve a single base, occlusion 64 bp, attention rollout 128 bp. Nothing is stretched into a resolution it does not have.
What is actually there, from the databases. Genes, CDS, introns, uORFs, tRNA and snoRNA genes, replication origins, LTRs and transposons come from SGD; regulatory regions from ORegAnno; binding sites from the Harbison 2004 / MacIsaac 2006 conserved regulatory code and from a JASPAR PWM scan. Provenance is drawn, not just recorded: a ChIP-supported site is solid, a conserved-only call is hollow, a PWM match is a hairline — three different strengths of evidence that must not be read as one. A grey cap marks a feature clipped by the window edge, so that edge is never mistaken for a real boundary.
Nothing here waits on a model. Every preset locus was run offline at the full 16,384 bp context and its predictions ship with the page. Loading the model — a 28.6 MB download and about 17 s of WebAssembly a run — buys the live layer activations, sequence editing and motif knockouts, not a curve that already exists. A dashed line, once you edit, is the unedited reference.
Why coverage is drawn at 16 bp, and what the flanks are for
Shorkie's head emits 896 bins of 16 bp, so 16 bp is the model's own resolution and no finer level exists. Drawing it per base would be storing 16 copies of one number and claiming a resolution the network does not have.
The flanks are real. Every output bin's receptive field is all 16,384 bases, so the 896-bin curve is drawn where it falls with its cropped flanks shaded, rather than stretched to fill the window. How much of that context the prediction actually depends on is measured in act 2.
The curve is a single forward pass, not rc-averaged, matching `make_predictions.py`. Every attribution below it IS rc-averaged, as every published Shorkie run is. The model is not rc-equivariant, so the two are deliberately different and saying so is better than quietly averaging one to match the other.
Every output track
- Question
- Is the curve above representative, or an artefact of averaging thousands of experiments?
- Why this matters
- The headline curve is a mean over thousands of experiments, and a mean can be shaped by a handful of loud tracks. If it is, every attribution built on it is explaining an artefact of averaging.
- Method
- All 5,215 predicted tracks as one row each, in the released targets sheet's order, with any single one plotted on the shared axis.
- Intuition
- A summary statistic cannot report on its own adequacy. Drawing all 5,215 rows shows structure within a block if the group mean is unrepresentative, and a uniform block if it is not.
- Cost
- Free. The 5,215 × 896 plane ships per locus, 2.15 MB.
- What would refute it
- If the four assay blocks did not read as bands the track ordering would be wrong — which it was once, and the curve labelled RNA-seq was mostly ChIP-exo.
What the answer tells us Where a block is uniform the group mean is a fair summary and every attribution built on it inherits that. Where it is not, the mean is the wrong object and the single-track view is the honest one.
One row per predicted experiment, in the order the released targets sheet lists them, so the four assay blocks read as bands. Click a row to plot that single track below — the overview above averages thousands of experiments, which is a curve no instrument would produce. Where rows outnumber pixels each drawn row is the maximum over the tracks it covers.
4Which bases drove it?
"Which bases mattered" has no single estimator. A derivative, a path integral, a finite difference and an ablation measure different mathematical objects, and a page that shows one of them shows a choice it has not justified. This act shows four over the region you trace and treats their disagreement as data: where they agree the model is locally linear, and where they part it is not.
Every map in this act has a spread, and it is measured. Recomputed across all 8 released training folds, 85% of the strongest bases keep their sign and a 95% interval clear of zero — against 54% of distance-matched controls. Reading the other strand moves an effect 0.943× as much as retraining does — comparably, and neither is noise. Act 9 has all eight checkpoints, the per-locus table and the gates.
Four attributions, and one that is not
- Question
- Which bases drove the prediction over the region you traced?
- Why this matters
- Attribution methods are routinely presented as interchangeable views of one truth. They estimate different mathematical objects, and treating a disagreement as noise hides that one of them may be answering a different question.
- Method
- Four estimators of the same scalar, all rc-averaged: a finite difference (mutagenesis), a local derivative (gradient × input), a path integral (integrated gradients) and a window ablation (occlusion).
- Intuition
- Any single map is a choice of estimator presented as a fact. Four that must coincide wherever the model is locally linear turn that choice into a measurement: their agreement bounds how much of the map belongs to the model rather than to the method, and their divergence localises where the model is doing something no first-order summary can carry.
- Cost
- Mutagenesis 23.8 ms a substitution → 19.5 min a locus; IG 32 forward+backward passes; occlusion 512 passes; gradient × input one backward pass.
- What would refute it
- They are not expected to agree. On TDH3's gene body grad × input against IG is r = 0.874, against occlusion 0.901, IG against occlusion 0.856 — and against mutagenesis per base, 0.658, which is the gap between a slope and a finite jump.
What the answer tells us Agreement between estimators of different mathematical type is evidence about the model; disagreement is evidence about where the model stops being locally linear. Neither is noise, which is why the page draws all four rather than averaging them.
The lanes themselves are in the viewport above — every method on one zoomable axis, becoming letters once a base is wide enough to be a letter. What is here is what the drawing cannot say: what each method measures, what it costs, and where they disagree.
Three of them draw letters, and they do not all draw the same number. A logo of the whole 16,384 bp window is not a drawing that exists — one base is 0.078 px — so letters appear only in a zoom. Mutagenesis ships every substitution, so it can stack up to four letters in a column; the checkbox above switches between that and the paper's projection onto the base that is actually there. Gradient × input and integrated gradients both multiply by the input, and the input is one-hot, so both are identically zero at the three bases that are not there and draw exactly one letter a column. That asymmetry is a fact about the methods, not a simplification of the drawing. Attention rollout resolves 128 bp and occlusion 64 bp, so neither becomes letters at all.
And they are not all scored on the same thing. Mutagenesis is logSED on this window's own gene body and needs no selection — the plane is a property of the locus. The other three are conditioned on whatever region you traced. So the mutagenesis lane routinely peaks over a different part of the axis than the lanes below it, and reading that as disagreement between methods is reading a difference in the question. Each lane names its own scoring target on its right.
Four methods, one axis, one annotation. They do not measure the same thing and are not expected to agree everywhere: mutagenesis substitutes one base and re-runs the model; gradient × input is a local linear sensitivity at the reference base, mean-centred across the four bases as Borzoi does; integrated gradients accumulates that gradient along a path from an all-zero baseline and is the only one here whose values sum to the prediction difference — a property it would lose if it were mean-centred, so it is not; occlusion ablates a whole 64 bp stretch and measures what the model loses. Where they agree the result is robust; where they diverge, the divergence is the finding. A method that covers only part of the window is drawn in place with the rest left blank, which is what the paper does for its own partial windows.
Traceback. Gradient × input: how much the prediction over the selected bins moves per unit of each input base, times the base that is actually there. It is a local sensitivity, not a decomposition — the values do not sum to the prediction. Dragging is exact rather than interpolated, because gradients superpose: the attribution for a set of bins is the sum of their individual gradients. Dragging resolves to 128 bp; picking a gene gives single-base resolution.
Two sources, and they are not the same method. Mutagenesis is the paper's Figure 4 quantity: every substitution actually run, then mean-centred across the four bases and projected on the reference, so the height reads as “how much the base that is there matters” rather than “what this one substitution did”. It now covers every base of the window — all 49,152 substitutions per locus, on both strands — so it is the primary track here and needs no caveat about where it stops. Gradient × input is a local linear sensitivity — it is identically zero at the three bases not present, because the input is one-hot, so it draws one letter per position. The paper's published figures use no gradients at all; its single gradient routine is unreached scaffolding that still declares a 4-channel input the 170-channel model cannot accept. Gradient × input is here because it is instant, not because it is what the paper did.
What the whole window shows that a promoter window cannot. Running all 16,384 bp on both strands for every locus makes one result visible that no single window does: the sign of the strongest substitution follows the gene's expression state in 22 of 23 windows. The three genes repressed in this baseline — HOP2 (meiosis-specific), GAL3 and GAL1 (galactose regulon, in glucose) — all have a largest substitution that raises predicted expression (+1.38, +1.12, +0.78 logSED, with 20–33 of their top 40 positive), and nineteen of the twenty expressed genes are negative. For a gene already being transcribed the model's biggest lever is breaking what drives it; for a silent one it is breaking what holds it down.
The exception is worth more than the rule. PWP1 is expressed here and its largest substitution is nonetheless +0.223. It is a ribosome-biogenesis gene of the RRB regulon, and Supplemental Figure S20B labels this very window with both of that regulon's repressor motifs — the PAC motif and RRPE. A positive lever on a gene carrying two repressor elements is mechanistically reasonable, but it was classified as expressed before the number was computed, so it stands as an exception to the rule as stated rather than as a confirmation of it.
Why that is not just an artefact of the log scale
The objection. logSED is a log ratio, and the repressed genes have the lowest baseline coverage — 77, 213 and 502 against up to 48,547. A given absolute change in predicted coverage moves logSED 161× more at GAL3 than at TDH3, so low-baseline genes should show large log effects whatever the biology.
The control, from the data itself. Baseline does not order the signs. The cleanest case is POP4: it has the second-lowest baseline of all twenty-three windows (166, below GAL3's 213 and GAL1's 502) and its largest substitution is −0.65, the most negative on the page — while PWP1, at nearly six times that baseline (958), comes out positive. DTD1 (299) and MMS2 (556) are likewise firmly negative between the two GAL genes. If the sign were a property of the log scale, the lowest baselines would carry the positives; they carry the largest negative instead.
The magnitude is a different matter, and the orderings invert. As a fraction of its own baseline the largest lever on the page is HOP2's +160%. In absolute predicted coverage HOP2's best substitution is worth +125 while PDC1's costs −8,841, seventy times more. Neither reading is wrong; they answer different questions. logSED is deliberately a ratio — that is what makes a silent promoter and a maximal one comparable — so "a stronger effect" is meaningless here without saying stronger relative to what.
What the old 500 bp window would have missed. It covers 3.1% of the sequence and holds a median 17.6% of the window's total |logSED| — genuinely well chosen, 5.7× enriched. But the strongest substitution falls outside it in 4 of the 14 windows that window ever covered — including HOP2's +1.38, the largest effect in the whole set. That comparison is scoped to those fourteen because the 500 bp window is the thing being compared against, and it never existed for the other nine.
How the four method tracks are computed
All four answer the same scalar, so they can be laid on one axis: f = log2( Σgene bins mean384 T0 tracks coverage + 1 ). That is the paper's own logSED target (ensemble.py:97-104): mean over tracks, sum over bins, +1 pseudocount, log base 2. Being a log ratio is what makes a silent promoter and a maximal one comparable.
- Gradient × input. One backward pass of
fw.r.t. the one-hot input gives [16384 × 4]. Mean-centre across the four bases (g − meanACGT g) — the Borzoi convention — then multiply by the input and sum the base axis. Because the input is one-hot, three of every four values are exactly zero, so this is one number per position. Cost: 1 backward pass. Reads as: if I nudged this base up a little, how would the prediction move. Fails when: the model is locally flat but globally sensitive — a saturated promoter can show near-zero gradient at a base that matters enormously. - Integrated gradients. The same gradient, accumulated over 32 steps along a straight path from an all-zero-DNA baseline (species channel kept, because a baseline with no species is not a sequence this model was trained to see), then multiplied by
x − baseline. Not mean-centred, deliberately: completeness — that the values sum tof(x) − f(baseline)— is a telescoping integral of the raw gradient, and centring destroys it. Measured, centring pushed the error from ~0.05 to 8–650%. Cost: 32 forward+backward per region. Reads as: this base's share of the whole difference between the real sequence and an empty one. Unique property: it is the only method here you can check against its own total, and the panel prints that check. - Attention rollout. The eight attention matrices, each mixed half-and-half with the identity to account for the residual stream, renormalised and composed:
∏ (½Aℓ + ½I). Row i then reads as where position i's representation came from. Cost: free — the maps already ship. Reads as: what the transformer can see for this region. Not an attribution: it says what could influence what, never what changed the prediction, and it is unsigned by construction. - Occlusion. Zero the four DNA channels across a 64 bp window and re-run; the change in
fper output bin is that window's effect. Cost: 256 real forward passes × 2 strands. Reads as: what the model loses without this stretch. Fails when: two stretches are redundant — remove either alone and nothing happens, which occlusion reports as "neither matters".
Both strands, averaged, for all four, matching Borzoi and the --rc every published Shorkie ISM run passes. Worth knowing what that is: this model was not trained with reverse-complement augmentation (augment_rc: false in all four params.json), so it is not rc-equivariant — measured on TDH3 the target reads 15.60 forward against 14.23 reversed and the two gradients correlate at 0.31. Averaging them is a test-time augmentation, not a free symmetry.
How much do they actually agree? Measured on TDH3's own gene body, at the 64 bp grid occlusion resolves: gradient × input against integrated gradients r = 0.874, gradient × input against occlusion r = 0.901, integrated gradients against occlusion r = 0.856. The three full-window methods agree strongly, which is the reassuring part. Against mutagenesis, per base over its own 500 bp, gradient × input falls to r = 0.658 — and that is expected rather than alarming: a gradient is a local linear sensitivity and a substitution is a finite jump to a different base. Where a promoter is saturated the two genuinely differ.
One thing the curve above does not share. The drawn coverage is the 3,053-track RNA-seq group mean; every attribution scores the 384 _T0_ subset the paper uses. Measured, the two correlate at r = 1.0000 and differ by 1% at the peak, so the distinction does not affect reading — but it is a distinction.
What full-window mutagenesis costs, and why it was once refused. 16,384 × 3 substitutions on both strands is 98,304 forward passes a locus and 1,376,256 for all fourteen. An earlier round priced that at 39.6 h and dropped it — but that was measured through onnxruntime on the CPU, on a graph whose batch axis is pinned at 1, so it baked in both the slowest engine available and the impossibility of batching. Neither limit belongs to the model. Re-measured on the same machine's GPU at batch 32, with the output head sliced to the 384 tracks actually scored: 10.47 ms a forward pass — and a substitution costs two of them, because every published Shorkie run averages both strands. Measured end to end: 23.8 ms a substitution, 19.5 minutes a locus — so the 23 windows the page now draws are about 7.5 h of GPU. (The 4.6 h figure recorded earlier was a timed benchmark over fourteen loci, which is where the rate comes from; it is not the size of the shipped set.) The engine change is not a precision compromise either — against the CPU the GPU agrees to 6.6 × 10⁻⁷ relative, three orders of magnitude tighter than the fp16 graph the earlier packs were built from.
The five lanes, side by side
- Question
- What is each method actually measuring, and where can it not be trusted?
- Why this matters
- A reader comparing lanes needs to know which differences are findings and which are properties of the estimator. Resolution, sign convention and scoring target all differ, and any of them can masquerade as disagreement.
- Method
- A side-by-side of mathematical nature, resolution, cost, completeness, mean-centring, sign and scoring target — every entry measured elsewhere on this page.
- Intuition
- Tabulating each estimator's own properties separates 'these methods disagree about this locus' from 'these methods cannot agree, because one resolves 64 bp and the other resolves one base'.
- Cost
- Free. Nothing here is a new run.
- What would refute it
- A row with a blank failure mode would be a method with no known failure mode, which does not exist.
What the answer tells us Read any divergence against this table before reading it as biology. Two lanes that cannot resolve the same feature are not in conflict; two that can and still differ are.
Four of the five differentiate the same scalar, which is what lets them share an axis. Attention rollout is the exception, and the table says so on every row: unsigned, summing to nothing, reporting what the transformer can read rather than what moved the prediction. It sits beside the other four because that contrast is the point, not because it is a fifth estimator. The four that do share the scalar differ in every other respect, and the differences are not stylistic — a method's resolution decides whether it can draw letters, its mean-centring decides whether its values sum to anything, and its scoring target decides which part of the window it peaks over.
| mutagenesis | gradient × input | integrated gradients | occlusion | attention rollout | |
|---|---|---|---|---|---|
| nature | finite difference | local derivative | path integral | window ablation | architectural reachability |
| resolution | 1 bp | 1 bp | 1 bp | 64 bp | 128 bp |
| cost per locus | 98,304 passes · 19.5 min | 1 backward pass | 32 fwd + 32 bwd | 512 passes · 22 s | free — the maps ship |
| letters a column | up to 4 | 1 | 1 | — | — |
| values sum to Δf | no | no | yes, to 0.14–9.41% | no | n/a |
| mean-centred | yes (Borzoi) | yes (Borzoi) | no — centring destroys completeness | no | n/a |
| signed | yes | yes | yes | yes | no, by construction |
| scored on | the window's own gene body | the traced region | the traced region | the traced region | the traced region |
| needs a region | no | yes | yes — and an exact anchor for single bases | yes | yes, and a live run |
| fails when | nothing — it is the ground truth of model behaviour, at 4,000× the cost | the model is locally flat but globally sensitive: a saturated promoter reads near zero at a base that matters enormously | the all-zero baseline sits above 62% of the genome, so the path runs downhill almost everywhere and the lane comes out 56.2% negative | two stretches are redundant — remove either alone and nothing happens, which occlusion reports as “neither matters” | always, as an attribution: it says what CAN influence what, never what changed the prediction |
The bottom row is the one to read first. Every method here has a failure mode, they are different failure modes, and that is precisely why the page draws five lanes instead of picking a winner. Where four of them agree the result is robust; where one dissents, its own row usually says why.
What drives what
- Question
- Which input stretches does which output region depend on?
- Why this matters
- Every other method here collapses the output axis to one scalar, which discards where in the gene an input stretch acts. Local and long-range effects then look identical.
- Method
- Zero the four DNA channels across a 64 bp window, re-run, and read all 896 output bins at once — the complete 256 × 896 matrix.
- Intuition
- One ablation answers all 896 output bins at once, so the complete input-region by output-region matrix costs no more than a single-scalar sweep. Two dimensions are affordable here and nowhere else on this page.
- Cost
- 256 ablation windows × 2 strands = 512 forward passes a locus. The only two-dimensional method here — one pass answers every output bin at once.
- What would refute it
- The diagonal dominates per cell — but it is one window in 256, so summed over a row the local footprint is only 0.2–5.6% of the most damaging window's total. Either half alone is the wrong answer.
What the answer tells us A narrow diagonal says the model is local, and the mass off it says how much of an effect travels. Reporting only one would make Shorkie look either purely local or purely long-range, and it is neither.
Every input window against every output bin. Each column is a 64 bp stretch of the input, ablated; each row is one of the 896 output bins; the colour is what that bin loses when that stretch is gone. Blue is a loss, red a gain. This is the only genuinely two-dimensional view on the page — every other method collapses to a single profile over the input — and it is measured, not inferred: 256 ablation windows on both strands, 512 real forward passes a locus.
Read it as a map, and read both parts. The dark diagonal is local effect — a window damaging the output directly above it — and per cell it is by far the strongest thing here. But it is also narrow: one window of 256. Summed across the row, the local footprint accounts for only 0.2–5.6% of the most damaging window's total effect, so the diagonal is where the model is most intense and the off-diagonal is where most of what it does actually lives. Neither statement alone is the truth. Off the diagonal, a vertical stripe is an input stretch many outputs depend on and a horizontal stripe is an output bin that reads widely.
The edges are the margins. The strip along the top is the column sum — how much each input window matters summed over every output — and the strip down the right is the row sum, how much each output bin depends on the sequence at all. Click a row to read that one output bin's input profile, or a column to see which outputs that window drives; click again to clear. Bins differ in expression by orders of magnitude, so the checkbox rescales each row to its own range when the raw map is dominated by the loudest ones.
Ablation zeroes the four DNA channels, which is exactly how the paper's language model masks a position and is indistinguishable from a run of N. That is a different question from the motif knockouts above, which shuffle: a shuffle keeps base composition and asks whether the arrangement matters, while zeroing asks whether the stretch carries information at all.
How the occlusion map is computed
For each of 256 windows of 64 bp: zero the four DNA channels across that window, run the model, and record log2(alt + 1) − log2(ref + 1) for every one of the 896 output bins. One forward pass answers a whole row, which is why a complete [256 × 896] map costs 256 passes — 27 s a locus, 54 s with both strands — and why this is the cheapest exact method here.
Per cell the unit is logSED, a log2 fold change, so bins of wildly different expression are comparable within a row. The default colour scale is shared across the whole matrix for that reason; the checkbox rescales each output bin to its own range, which trades the between-bin comparison for the within-bin one and is worth it when a few loud bins dominate.
The margins are plain sums of |logSED|: the top strip is each input window summed over all 896 bins, the right strip each output bin summed over all 256 windows.
The diagonal is not the identity line. The two axes cover different spans — the input is the whole 16,384 bp window, the output only the cropped 1,024–15,360 interior — so it is a line of slope <1 that reaches neither corner. It is drawn dashed where it actually falls.
Zeroing is not shuffling. Zeroing the DNA channels is how the paper's language model masks a position and is indistinguishable from a run of N: it asks whether the stretch carries information at all. The motif knockouts above shuffle, which preserves base composition and asks whether the arrangement matters. Different questions, and they can disagree.
What occlusion cannot see: redundancy. If two stretches carry the same information, removing either alone changes nothing, and this map reports both as unimportant.
5Is the attribution real?
An attribution map always looks like something. Nothing in act 4 could have come out empty, so nothing in act 4 is yet evidence that the model found anything real. These are the tests that can come out negative: attribution scored against annotation the model was never shown, every curated site knocked out and measured, and both re-checked on all twenty-three windows rather than the one on screen.
Does the attribution land on real biology?
- Question
- Does the attribution land on features the model was never shown?
- Why this matters
- An attribution map is a claim about the model, not about biology. Without an external reference it cannot be distinguished from a smooth function of position, and every peak looks meaningful.
- Method
- Mean |attribution| inside a feature class over the mean across the window, against a null of 256 deterministic circular shifts of the annotation — which preserves feature count, length and gap structure and destroys only alignment.
- Intuition
- The annotation was never shown to the model, so alignment with it is evidence the model found something real. The null must destroy alignment and nothing else — a circular shift keeps the feature count, every length and every gap, so it cannot be beaten by geometry alone.
- Cost
- Free. Both the annotation and the packs ship.
- What would refute it
- A class at 1.00× says the model puts no more weight there than anywhere else. Every class it can measure is shown, not the ones that came out well.
What the answer tells us Alignment with an annotation the model never saw is the first evidence here that an attribution is about sequence function rather than about the estimator. It remains an association under one null, not a demonstration that the model uses the feature.
One number per method, per annotation class. For each class, the mean |attribution| on the bases that class covers, divided by the mean across the whole window — so 1.00× means no preference. The null comes from circularly shifting the annotation against the signal, which keeps the number of features, their lengths and their spacing exactly and destroys only their alignment; the shaded figure beside each ratio is that null's spread. Cells are shaded by how far the ratio sits from its own null, not by its size. The table measures every evidence tier whatever the lanes above are drawing — comparing the three tiers against each other is the point, and a drawing toggle should not silently change a statistic.
This table is computed from one checkpoint. The attribution underneath it is gradient × input on fold f0, whose per-base signal survives retraining for 85% of the strongest bases and 54% of matched controls — act 9 measures it. An enrichment ratio inherits that: it is a statement about where this training run puts its weight.
Read the negatives too. A class at 1.0× is a real answer: it says the model puts no more weight there than anywhere else. And a method cannot resolve a feature finer than its own resolution — occlusion moves in 64 bp steps and attention rollout in 128 bp, so neither can say anything about a 7 bp binding site, which is why their columns are marked. This is a correlation against a stated null, not evidence the model uses the feature: a site that happens to sit inside a region the model attends to for unrelated reasons scores exactly the same.
How the enrichment and its null are computed
The statistic. Take the per-base signal s over the 16,384 bp window and a binary mask m for the class. The ratio is (Σ|s|·m / Σm) / (Σ|s| / L) — mean magnitude inside over mean magnitude everywhere. It is scale-free, so two methods with wildly different units are directly comparable, and it is a magnitude, so a class the model suppresses still reads as enriched. The mean signed value inside is reported alongside so that direction is not lost.
The null is a circular shift, not a resample. The mask is rotated by each of 256 evenly spaced offsets and the ratio recomputed. Rotation preserves the feature count, every feature's length, and the spacing between them — so the null compares against an annotation that looks exactly like the real one but is not aligned to the signal. Resampling positions instead would compare against a feature set that does not resemble the real one at all, and would call almost everything significant. The offsets are deterministic (evenly spaced, zero excluded), so a published number is reproducible rather than a draw.
The p is empirical, (#{null ≥ observed} + 1) / (256 + 1), so it is never exactly zero and its floor is 1/257 ≈ 0.0039. A cell at that floor means "no shift out of 256 reached this", not "p = 0".
What it cannot tell you. Circular shifts are a weak null against broad positional trends: if a method's signal is concentrated near the window centre and a class happens to be too, the shift will not fully break that. It is also a per-class test with no multiple-testing correction across the table — read a single marginal cell with that in mind.
Every curated site, knocked out
- Question
- Is an attribution peak a real motif, or base composition?
- Why this matters
- A peak on a binding site could be its arrangement or merely its base composition — poly-A tracts and GC-rich stretches draw attention for reasons that are not regulatory.
- Method
- Permute the exact bases inside each curated site — same length, same composition, only the arrangement destroyed — and measure the change over the window's own gene body, as a mean over 5 shuffles with its spread.
- Intuition
- Permuting bases inside the site holds length and composition exactly and destroys only the order. An effect that survives the permutation was never about the motif.
- Cost
- 5 shuffles a site, two forward passes each — forward and reverse complement, averaged per bin before the gene body is summed — around 50 sites a locus. Same rc-averaging as every attribution on this page, so the numbers are directly comparable.
- What would refute it
- A near-zero result is an answer about that site. A low-complexity site cannot be shuffled at all, so the sweep records how many DISTINCT permutations it achieved rather than reporting a zero spread as a measurement.
What the answer tells us An effect that survives a composition-preserving permutation was never about the motif's arrangement. A near-zero result is therefore an answer about that site rather than a failed measurement — provided the site had distinct permutations to draw from at all.
Not six hand-picked motifs — all of them. Every ChIP-supported binding site in this window was shuffled and the model re-run, five times per site, and what is reported is the mean logSED with its spread. One shuffle is one sample; a single draw presented as a measurement is exactly what this replaces. Scored over the window's own gene body on the paper's 384 T₀ tracks, both strands — the same quantity the interactive knockout reports, using the same seeded shuffle, so a swept value and a clicked one agree exactly rather than approximately.
Where a gene has a well-known regulator, the sweep tends to find it, unprompted. Ranked by the size of the effect, the strongest site is RAP1 at the glycolytic genes TDH3 and PDC1 and TYE7 — also a glycolytic activator — at PGK1; FHL1 at the ribosomal protein gene RPL26A; and a GAL-regulon factor at both galactose genes, GAL80 at GAL1 and GAL4 at GAL3. At GAL1 the effect is unusually clean: the six largest sites in the window are all GAL4 or GAL80, the activator and the repressor of that regulon. Nothing in this pipeline knows which factor is supposed to matter — the sites come from a database and the ranking comes from the model.
It is a tendency and not a law, and the table is ranked by magnitude so you can check it: at HOP2 and ACT1 the winner is STE12, which is a mating and filamentation factor rather than the regulator either gene is known for, and at KRE33 and DTD1 the winners are not characterised regulators of those genes at all. Sign is not reliable either — knocking out the GAL80 repressor site raises the prediction, which is the right direction for a repressor, but several activator sites also come out positive
A site whose bar is inside its own spread did nothing measurable, and that is a result about that site rather than a broken button. Shuffling preserves base composition and destroys only the arrangement, so this asks whether the order of those bases matters — a different question from the occlusion map, which zeroes a stretch and asks whether it carries information at all. The two can disagree. A low-complexity site cannot be shuffled at all — one site across the 23 windows is five Cs and an A, which has almost no distinct permutations — so those are labelled rather than reported as a null result. “Unshuffleable” and “the model ignores it” are different findings that produce the same number.
How the sweep is computed, and why a mean over shuffles
Per site: Fisher–Yates shuffle the bases inside the site, re-run the model, and record log2(Σ coverage + 1) over the gene body minus the same quantity on the unedited sequence. Repeat with k = 5 independent permutations and report mean ± sd. Cost is two forward passes per shuffle — the sequence and its reverse complement, 104 ms each through onnxruntime — so the whole ChIP-supported tier across all 23 windows is several thousand passes.
Why not one shuffle. A single permutation is one draw from the distribution of "sequences with this composition". Two shuffles of the same site can differ by more than the effect being reported, so the single number the page used to show was a measurement with an unstated error bar of unknown size. The sd here is that error bar, measured.
Why the gene body, not the peak. A 16 kb yeast window holds a dozen genes and the tallest is rarely the one whose promoter was edited. Measuring globally reports a number about an unrelated gene — on one window the global peak is 114.3 while the named gene's own body peaks at 7.8.
What it cannot tell you. Shuffling a site does not remove the possibility that the model reads a different, overlapping element; sites here overlap frequently. And a null effect at a real binding site can mean the model is robust to it, that a redundant site elsewhere covers it, or that the site is not used in the condition the T₀ tracks represent.
What all twenty-three windows say
- Question
- Do the single-locus results hold across the whole set?
- Why this matters
- The page's strongest claims — that evidence tier predicts attribution, that introns are enriched — would be anecdotes if they rested on whichever window happens to be on screen.
- Method
- Every enrichment and every knockout re-measured on all 23 windows and reported as medians, plus attribution aligned on the direction of transcription across 198 genes.
- Intuition
- A median across every window is robust to one locus behaving oddly, and reporting the spread beside it shows whether a claim is a tendency or something that holds everywhere.
- Cost
- Free. Derived from the shipped packs.
- What would refute it
- A median over 23 can contradict any single window, and does: the 3.26× ChIP-supported enrichment quoted from one locus is 1.73× across the set.
What the answer tells us The medians here are the page's cross-locus claims; any single-locus figure beside them is an illustration. Where the two disagree, the median is the claim and the locus is the anecdote.
Every other panel in this act shows one window at a time. This one is the pattern across all twenty-three, computed once from the same packs and the same null those panels use — so it cannot drift from them.
Each row is one annotation class; the bar is the median across the windows that contain it, and the dots are the individual windows. A class present in two windows and a class present in 23 are drawn the same width, so the count beside each row is what says how much to trust it.
How the cross-locus summary and the TSS profile are computed
Nothing here is new data. The per-class enrichment is the same statistic the enrichment panel above computes, run once per window offline and the medians taken; the knockout winners are read from the same -ko.json packs the sweep panel reads. It is a second view of numbers that already exist, not a second copy of them — verify_pipeline re-derives it from the packs and fails if the two disagree.
The TSS profile aligns on the direction of transcription, not on coordinate order: the start is txStart on the plus strand and txEnd on the minus, and minus-strand profiles are reversed so "upstream" is the same side of the plot for every gene. Without that flip the average puts promoters against terminators and produces a flat curve that looks exactly like a real null result.
Each gene is normalised to its own total before averaging, so a highly expressed gene does not dominate. The band is the spread across genes, not a confidence interval on the mean — these genes differ by orders of magnitude in expression and the spread is the honest statement of that.
What it cannot tell you. 198 genes from twenty-three windows is a small, non-random sample: they were chosen as classic high expressers, the galactose regulon and the paper's own figure panels. A metaprofile from them describes those genes, not the yeast genome.
6How does it do it?
An attribution says which bases mattered and nothing about how. Between the sequence and the prediction sit twenty stages, and whether the answer is decided early, at the bottleneck, or on the way back out is a different question with a different kind of evidence. This act opens the model in six steps that go progressively deeper into the network's own basis. Watch a sequence pass through all twenty stages; open any one of them; trace a region back through every layer to the bases it came from; ask what the individual attention heads read; then ask whether the network's channels are the right unit at all, and finally at what depth the answer is already decided.
The network, end to end — watch a sequence pass through it
- Question
- What shape is the computation between the sequence and the prediction?
- Why this matters
- The architecture is usually described in prose, where 'a U-Net over a transformer' conveys nothing about the sizes involved. Which stage compresses what is the fact every internal panel below depends on.
- Method
- All 20 stages at their true tensor sizes — height is positions, width is channels, both log-scaled over the range actually present so the U is visible.
- Intuition
- Drawing each stage at its true tensor size on log axes makes the U visible and makes a skip horizontal — a skip joins two stages of equal resolution, which is exactly why it can carry information around the bottleneck.
- Cost
- Free. The per-stage maps come from the same forward pass as the prediction.
- What would refute it
- Not a measurement: this is the architecture, checked against the checkpoint's own tensor shapes.
What the answer tells us The shape decides which later panels are even possible: a bottleneck at 128 positions is why attention resolves 128 bp, and the three skips are why a bottleneck-only account of this model is incomplete before any measurement is taken.
Each block is one stage, sized by its real dimensions on the log scales named above. Click a stage to inspect it. Before you run the model the blocks are empty outlines; afterwards they fill with the activations from that same forward pass.
Switch to “relevance” and the same blocks show one region instead of the whole window. Trace a region below, and every stage repaints with what it contributed to that region alone — every other output bin masked out. Both of the map's margins are exact: the per-channel one and the per-position one are each a row-sum of precomputed groups, because gradients are linear in which outputs you select. Its interior is their outer product, which is the only reconstruction consistent with both margins if channel and position act independently — a real assumption, and the reason the panel names it.
This layer, in detail
- Question
- What is one stage actually computing?
- Why this matters
- A stage summary — a norm, a mean — cannot tell a layer where one channel does everything from one where the work is spread over hundreds. Those are different mechanisms with the same summary.
- Method
- Every channel of the selected stage as one row, on a diverging scale centred at zero, normalised by percentile rather than by min–max.
- Intuition
- A residual stream is signed, so 'idle' and 'strongly negative' are different states that a one-sided ramp draws identically — reading sign as activity once made 90% of cells look busy on most stages. Centring at zero is what lets an idle neuron leave its cell empty, and percentile bounds keep a handful of outliers from setting the range for everything else.
- Cost
- Free. Precomputed per locus.
- What would refute it
- A uniformly inked raster encodes sign rather than activity — which the p1–p99 ramp did, putting 90–96% of cells above the ink floor on 15 of the 20 stages.
What the answer tells us A stage where a few channels carry the signal is one a sparse decomposition might improve on; a stage where it is spread evenly is one where per-channel readings mean little. The raster says which kind each stage is.
Trace a region — layer by layer, back to the sequence
- Question
- For the region you traced: which stages, which neurons, and back to which bases?
- Why this matters
- Knowing that a region matters does not say how. Without a path from the region back through the stages to the sequence, an internal panel is a picture of activity rather than an account of the computation.
- Method
- Per-channel and per-position relevance margins, both exact and both superposing, with the interior reconstructed as their outer product under independence.
- Intuition
- Both margins are exact and superpose, because gradients are linear in which outputs you select — so an arbitrary region needs no new model run. Only the interior is reconstructed, and it says so.
- Cost
- Free for any region — gradients are linear in which outputs you select, so an arbitrary contiguous region is exact with no model run.
- What would refute it
- The earlier estimate weighted the channel margin by activation and is essentially uncorrelated with the exact answer (r −0.08 to 0.23, argmax positions disagreeing entirely). The reconstruction is checked against its own margins rather than assumed.
What the answer tells us The exact margins say which stages and channels carry a region; the reconstructed interior says roughly where within them. Reading the interior as a measured neuron-by-position map would claim more than the computation supports.
Five questions, in order. Each one narrows the last: which stages carry this region, then which neurons inside the strongest of them, where those neurons fire, what they respond to in the curated annotation, and finally what the transformer could have read regardless of what moved the prediction.
1 Which stages carry this region?
Longer bar, more of this region's relevance passes through that stage. Click a stage to open it in This layer, in detail above; click one of its channel chips to plot that neuron in step 2.
Each bar is a mean over the stage's channels, never a sum: summing would rank a 1,536-channel stage above a 384-channel one on width alone, which is a fact about the architecture and not about this region. Stages whose activations live on their own tensors report own tensor rather than zero — "not measured here" is not the same claim as "contributes nothing".
2 Which neurons, exactly?
The top eight channels of the selected stage, each drawn as its real activation profile — nothing reconstructed. Enrichment is the share of a channel's activity inside the traced region over the share of the window that region occupies, so 1.0× means it fires there no more than anywhere else and the interesting rows are well above it.
Each trace is scaled to its own range, so heights are not comparable between rows — the fires pair gives the real range. A faint rule marks zero, drawn only where the channel actually crosses it; several fire entirely negative. These traces are the un-collapsed answer to the same question the relevance map asks in This layer, in detail above: that map reconstructs a stage's [channels × positions] relevance as the outer product of two exact margins, so its margins are measured and its interior is an independence assumption. These are measured throughout. Where the two disagree, trust these.
3 Where do those neurons fire?
One row per stage, the window along x. Red where that stage draws more than its own average, blue less, neutral where it has no positional preference. Read it top to bottom: structure appears only from the transformer on.
The early residual blocks come out close to neutral, and that is a real property rather than a flat drawing: pooled to 128 positions, a convolution over 16,384 bp is spatially almost uniform. Structure appears where the receptive field becomes the whole window and the model can choose where to look. Each row is the factorised combination — per-channel relevance times per-position activation — not a per-position gradient, because the shipped pack carries relevance summed over position. It answers “where do the channels that matter fire”, which is not quite “which positions the gradient flows through”.
4 What are those neurons responding to?
The same neurons as step 2, scored against the curated annotation instead of against the region: for each feature class, which of this stage's relevant channels fire on it more than chance. This is what turns “channel #14 matters here” into “channel #14 responds to binding sites”.
Scored on the 32 channels with the highest exact relevance for this region, not all of them — the question is what the neurons that matter here respond to. Both the activation profile and the annotation are pooled to the packs' common 128 positions, so one cell is 128 bp: a channel cannot be shown to respond to a 7 bp site, only to the 128 bp neighbourhood containing it. That resolution limit is why a high ratio here is a lead, not a conclusion.
How a neuron is matched to a feature class
Same statistic as the enrichment table, applied one channel at a time. The channel's real activation profile [128] plays the role of the signal; the annotation mask, pooled from [16384] to [128] as the covered fraction of each cell, plays the role of the weight. The ratio is mean |activation| where the class is, over mean |activation| everywhere, and the null is the same deterministic circular shift.
Pooling by fraction, not by max. A 7 bp site covers 5% of a 128 bp cell. Taking a max would mark that whole cell as annotated and make every class look identical once pooled, which is the failure that would make this panel meaningless while still producing numbers.
What it is not. A channel that responds to a class is not a detector for it: the pooled resolution is 128 bp, the classes overlap each other, and nothing here controls for a channel simply firing where the gene is. Read a row as "worth looking at", then check the neuron's own trace in step 2.
5 What could the transformer read at all?
Attention rollout: the eight attention matrices composed, each mixed half-and-half with the identity to account for the residual stream. This is architecture, not attribution — it says what information could reach this region, never what changed the prediction, and it is unsigned by construction.
It is a second, independent answer, and it needs no new data: the eight [128 × 128] maps already ship in every pack. Where it disagrees with gradient × input, the disagreement is the finding — a region the transformer can see but does not use is a different statement from one it cannot see.
How the layer-by-layer trace is computed, and what each collapse throws away
Every view in this panel starts from the same tensor and collapses it differently. For one stage, relevance = |∂f/∂a ⊙ a| over its activations a: [channels × positions] — how much each neuron, at each place, moved the prediction over the selected bins. That array is far too large to ship for every region, so what is precomputed is its two margins:
- Channel margin — sum over position, giving [channels]. "Which neurons mattered, wherever they fired."
- Position margin — sum over channel, then sum-pooled to a common 128, giving [128]. "Where in the window this stage drew from, whichever neuron did it."
Both are exact, and both superpose: gradients are linear in which outputs you select, so an arbitrary contiguous region is the row-sum of the precomputed 112 groups of 8 bins — no model run, no interpolation. That is why dragging is instant and still exact.
The map's interior is not exact. Reconstructing [channels × positions] from two margins requires an assumption, and the one used is independence: the outer product, normalised to sum to 1. It is the unique reconstruction consistent with both margins, and it is still an assumption. The neuron traces below the map exist precisely because they are not: each is one channel's real activation, straight from the shipped stage maps.
The stage stack is the position margin, one row per stage, each row scaled by its own 99th percentile and then painted against its own mean — red above average, blue below, neutral where the stage has no positional preference. Per-row scaling because stages differ by orders of magnitude; against the mean because the early residual blocks are genuinely near-uniform once pooled to 128, and a from-zero ramp draws that truth as a saturated bar reading "maximally relevant everywhere".
Three stages are missing from all of this, and that is not an oversight: the input one-hot, the conv stem and the output head live on their own tensors in the exported graph and have no per-layer relevance in the pack. They show their activations and say so.
The enrichment number on each neuron trace is the share of that channel's Σ|activation| falling inside the traced region, divided by the share of the window the region occupies. 1.0× means the neuron fires there no more than anywhere else. Measured on block 4, the top-relevance channels come out 0.85–1.32× — so a channel's relevance to a region generally comes from what it computes there, not from firing only there.
Which attention layer reads what?
- Question
- Do the transformer's attention layers specialise on kinds of sequence?
- Why this matters
- Attention maps are the most over-read object in interpretability. Establishing what this one can and cannot support is worth more than another visualisation of it.
- Method
- Each layer's attention mass — its four heads averaged, which is how the pack ships — against annotation classes, with the mask pooled by MEAN (a 7 bp site is 5% of a 128 bp cell) and a circular-shift null.
- Intuition
- Attention is a distribution over keys, so a row's mass on annotated positions asks what fraction of what a layer reads is annotated — a quantity a merely loud layer cannot inflate.
- Cost
- Free. The 8 × 128 × 128 maps ship in every pack — no forward passes at all.
- What would refute it
- Every class records its ceiling, 1/coverage. Genes cover 91.5% of a window so nothing can exceed 1.09× on them, and reading their 1.03 as flat would be wrong.
What the answer tells us Enrichment says where a layer's attention mass falls under one null. It does not say the layer uses what it attends to, and with the four heads averaged before export it cannot speak about specialisation within a layer at all.
The transformer at the bottleneck is eight layers of four heads over 128 positions, and every locus pack already ships their maps — so this asks a mechanistic question with no forward passes at all. For each layer, how much of its attention lands on each class of curated annotation, against a null that rotates the annotation and destroys only its alignment.
Layers, not heads, and the distinction is forced by the pack. The export averages the four heads within each layer before writing the map — attention.mean(dim=2) in build_onnx.py — so what ships is eight layer matrices, not thirty-two head matrices. Every number below is therefore a statement about a layer with its heads already collapsed. Per-head specialisation is a different and finer question, and answering it would need a pack this page does not have. Not a live run either: the export averages the heads before the graph is written, and the browser consumes that collapsed output, so recovering the four heads needs a newly instrumented pass against the raw checkpoint — four times the attention payload. This panel does not answer it, and an earlier version of this text claimed it did.
The layers place their attention mass the way a biologist would divide the genome. Layers 0 and 3 are enriched on regulatory DNA against the shifted null — conserved binding sites at 1.75× and 1.88×, ORegAnno regions at 1.82× and 1.72×, ChIP-supported sites at 1.65× and 1.52× — and depleted on coding sequence at 0.73 and 0.69. Layer 4 is the mirror image: CDS 1.10× and every regulatory class below 0.90. Layer 6 alone is enriched on tRNA genes, at 1.50×. Nothing in training named any of these categories. Enrichment is where attention mass lands, not evidence that the layer uses it: this is a descriptive alignment under one null, and it is unsigned token mixing rather than attribution. And because the heads are averaged, a layer reading “flat” here could still hold one specialised head and three that cancel it.
The three things that decide whether this means anything
The mask is pooled by MEAN, never by max. One bottleneck position is 128 bp and a typical binding site is 7 bp — 5% of a cell. A max would mark the whole cell annotated and every class would come out identical once pooled: numbers that mean nothing while looking exactly like numbers that do.
The null is a circular shift, not a resample. Rotation preserves the feature count, every length and every gap, and destroys only the alignment. Resampling positions compares against an annotation that does not resemble the real one and calls almost everything significant.
Every class carries its ceiling, 1/coverage. Genes cover 91.5% of a window, so no layer can exceed 1.09× on them however hard it looks — reading their 1.03 as "flat" would be wrong when it is most of the way to the maximum the geometry allows. The same saturation makes phastCons look uncorrelated inside coding sequence.
Is the model's own basis the right one? sparse dictionary · 6,144 features
- Question
- The layer panel draws one row per channel and the traceback ranks channels by relevance. Both assume a channel means something. Does it?
- Why this matters
- Two panels above rank channels and draw one row per channel. Both assume a channel is a meaningful unit; if the model's features are spread across channels, those panels are reading an arbitrary basis.
- Method
- A TopK sparse autoencoder on the attn_out8 residual stream: 384 dimensions expanded to 6,144 features with exactly k = 32 active, decoder columns unit-norm, trained on 94,970 bottleneck vectors swept from the genome.
- Intuition
- If a sparse overcomplete dictionary reconstructs the same activations far better than the identical architecture trained on column-shuffled activations, the structure it found was in the co-activation and not in the marginals.
- Cost
- One forward pass a genome window to collect, then 4,000 Adam steps. Minutes, not hours — the bottleneck is only 384 wide.
- What would refute it
- Two controls, both run. Training on column-shuffled activations keeps every channel's marginal distribution and destroys only the co-activation structure: it reconstructs at FVU 0.373 against 0.019. And the 384 raw channels are scored against the annotation on identical terms — if they grounded as well, the dictionary bought nothing. They reach 4.715x against the features' 5.977x.
What the answer tells us If the dictionary beats a composition-matched control, the structure is in the co-activation rather than in the model's own channel basis — and the two panels above, which rank and draw channels, are reading a basis rather than a set of features.
A channel is a direction the optimiser happened to land on. In a network carrying more concepts than it has channels, the concepts sit in superposition — spread across non-orthogonal directions, so one channel fires for a binding site and a splice donor with nothing connecting them. A sparse dictionary looks for a better basis: overcomplete, so there is room for one direction per concept, and hard-sparse, so each direction is under pressure to mean one thing.
| real activations | column-shuffled control | |
|---|---|---|
| fraction of variance unexplained | 0.0194 | 0.3733 |
| features alive | 6,144 of 6,144 | 6,144 of 6,144 |
| mean active features a position | 32.0 | 32.0 |
The reconstruction gap is the first finding. With 32 of 6,144 features active at a time the dictionary reconstructs the bottleneck to within 1.9% of its variance. The same architecture, sparsity and step count on column-shuffled activations reaches only 37.3% — a gap of 35.4 points. Shuffling each channel independently keeps every channel's own distribution and destroys only which channels fire together, so that gap is exactly the co-activation structure the model's own basis was spreading out.
Does a feature mean anything?
Reconstructing well is necessary and not sufficient. Every feature is scored against the genome-wide annotation with the same circular-shift null the enrichment table uses, and the 384 raw channels are put through the identical test — same number of candidates, same null — because if the basis the dictionary replaced grounds just as well, the dictionary bought nothing.
| feature | fires on | strongest class | enrichment | most enriched 6-mers |
|---|---|---|---|---|
#5309 | 847 cells | tRNA gene | 174.7× · p 0.0039 | CCCGCG (31.52×), CGAACC (29.71×) |
#894 | 829 cells | tRNA gene | 169.8× · p 0.0039 | CCGGGG (20.29×), GCCCCC (20.15×) |
#4464 | 1,838 cells | snRNA gene | 110.6× · p 0.0039 | CCGTAC (7.35×), ATCGCA (6.37×) |
#2557 | 2,215 cells | tRNA gene | 90.3× · p 0.0039 | GCTAAC (6.45×), TAGCTA (6.15×) |
#1301 | 1,502 cells | LTR | 76.9× · p 0.0039 | GGAATC (16.05×), CTAGGG (13.62×) |
#6020 | 1,563 cells | LTR | 64.5× · p 0.0039 | CTAGGG (10.89×), AACCCG (10.66×) |
The dictionary grounds better than the basis it replaced. Matched at 384 candidates each and the same null, the sparse features reach a median best enrichment of 5.977× against the raw channels' 4.715×, and a maximum of 204.042× against 13.642×. The top features spread over 19 annotation classes rather than piling onto one.
Read the classes before the ratios. The largest numbers land on tRNA, LTR and snRNA genes — classes with distinctive base composition, which a feature can find without learning anything about regulation, and the 6-mer column shows exactly that: the strongest tRNA features are GC-rich where the genome is not. The harder claim is the quieter one: 5 of the interpreted features ground best on a binding-site tier, the strongest at 9.7× on ChIP-supported sites. That is a real result at a much lower ratio, and it is the one worth trusting least casually.
Why a k-mer signature and not a position weight matrix
A feature fires at 128 bp resolution. A PWM built from 128 bp windows is near-uniform however it is built, and would still match a database at high Pearson because correlation normalises amplitude away — that is precisely the failure the first motif-discovery run shipped, at 0.00 bits. Counting 6-mers says what a cell contains without claiming to know where, which is the honest summary at this resolution.
The enrichment is computed in Python and must match the browser's. The signal is 94,970 positions × 6,144 features, far past what a page can compute, so the statistic is reimplemented in the generator. Its circular shifts divide by 257, which is prime and divides neither array length in use, so every offset reduces to an odd denominator and can never be exactly one half — there is nothing for JavaScript's round-half-up and Python's round-half-to-even to disagree about. The generator asserts this rather than inheriting it from the length it was first checked at.
What is still not claimed. A feature enriched on a class is a correlation against a stated null, not evidence the model uses that class: a feature firing on GC-rich sequence will score on tRNA whether or not it encodes anything about transcription. Nothing here has been causally tested by clamping a feature and re-running the model, which is the experiment that would settle it.
At what depth is the answer already decided? causal tracing · 22 windows
- Question
- Where in the network — at what depth, and at which positions — does the promoter's information become sufficient to reconstruct the prediction?
- Why this matters
- Attribution says which bases matter, not where in the network their information becomes sufficient. In a skip-connected model 'the bottleneck decides it' and 'the skips carry it' predict different things.
- Method
- Causal tracing (Meng et al. 2022). Corrupt the promoter by dinucleotide shuffle, then restore the CLEAN activations at one stage and one 512 bp band, and measure how much of the clean prediction returns.
- Intuition
- Restoring a whole stage is degenerate — everything downstream simply inherits the clean run. Restoring one spatial band at one stage is what makes depth and position separable at all.
- Cost
- 19 stages × 32 bands × 22 windows — 608 forward passes a locus, 346 s for the set.
- What would refute it
- Three controls, asserted per locus. Restoring every position of the first stage must recover exactly 1.0, restoring every position of the last stage must too by an independent route, and restoring nothing must recover exactly 0.0. Together they catch a patch written into the channel axis, which broadcasts cleanly and produces entirely plausible numbers.
What the answer tells us Where recovery concentrates is where the promoter's information becomes sufficient. A bottleneck that does not fully recover is not a defect but the skips measured, and it bounds how much of this model any bottleneck-only story can explain.
Every attribution above answers which bases mattered. None of them answers where in the network that mattering became decisive — and those are different questions, because a base can matter enormously while the representation carrying it only forms several layers in.
Why 22 windows and not 23. Recovery is a fraction of the gap the corruption opened, so a window whose gap is ~0 has no denominator. HOP2's promoter shuffle moves the prediction by less than 0.001 — it is meiosis-specific and silent in this baseline, so scrambling its promoter changes essentially nothing — and the generator skips it rather than dividing by noise and reporting a recovery of forty. An exclusion, not a missing measurement.
Recovery rises through the encoder and decays through the transformer. Averaged over the 22 windows it climbs from 0.434 at the stem to 0.546 at the last residual block, is essentially tied there with the first attention layer (0.543), and then falls monotonically across the remaining seven to 0.349. Per window the peak cell sits in the encoder in 18 of 22, at the bottleneck in 2 and in the decoder in 2 — and it is nearly total, a median of 96.5%.
So by the time a promoter reaches the last residual block, one 512 bp band of its representation is already almost the whole answer — and every attention layer after that makes a single band less sufficient, not more. Restoring one band at the bottleneck recovers a median of only 3.9%. That is what self-attention is for: the information is not lost, it is no longer in one place. Read the two claims together — the encoder localises and the transformer distributes — rather than either alone, and note that the encoder-versus-bottleneck margin at the peak is 0.003, which is a tie and is drawn as one.
The three controls, and what each one pins. Restoring every position of the first stage makes everything downstream the clean run's, so it must recover exactly 1.0 — and does, at every window. Restoring every position of the last stage must too, by a completely different route: the head is a deterministic function of it. Restoring nothing must recover exactly 0.0. Together they fix both ends of the scale and catch the one failure that otherwise passes silently — a patch written into the channel axis rather than the position axis, which broadcasts cleanly, completes without error, and produces entirely plausible numbers.
The fourth check, which is not a control — and what it turned out to measure
The obvious fourth control is not one. Restoring every position of a bottleneck stage should make everything downstream the clean run's, so it ought to recover exactly 1.0. The first version of this script asserted that and failed at 0.9663.
It is not a wiring bug. The U-Net skip connections carry the seven residual blocks straight to the decoder, around the transformer entirely — so a perfectly clean residual stream still meets corrupted skips. The shortfall is a measurement of how much of the prediction bypasses the bottleneck, and across these windows the median is 2.8%. Nothing else on this page reports that number, so it is published rather than worked around.
Why the corruption is a shuffle. Zeroing the DNA channels is indistinguishable from a run of N, so the corrupted run would be out-of-distribution rather than merely uninformative and every recovery figure would be measuring the model's response to impossible input. The shuffle is Altschul–Erikson and preserves every dinucleotide count exactly.
Why bands and not layers. Patching a whole stage makes everything downstream the clean run's, so recovery is 1 at every depth and the plot says nothing. Restoring a band is what makes the question answerable — and the two ends of that argument are the controls above.
7Grammar — what a first-order method cannot ask
The knockout sweep in act 5 is destructive: it shuffles a site already sitting in a real promoter and reports what the prediction loses. That measures whether a site is necessary there — and a site can be necessary for reasons that are not about the motif at all, because of the nucleosome-free stretch it sits in or the TATA box beside it. These six ask the opposite question constructively: start from shuffled DNA containing nothing, add one element at a time, and see what the model does. The Hessian panel relaxes the additivity every method above assumes; the last one turns the question round entirely and asks the model's own mutagenesis what a motif is, without a database to check against.
Is the motif enough on its own?
- Question
- Is a motif enough on its own to move the prediction?
- Why this matters
- The knockout sweep can only report necessity in the context a motif already sits in. A site can be necessary because of what surrounds it and carry nothing on its own.
- Method
- Implant the consensus into 200 dinucleotide-shuffled backgrounds and compare with its own scramble in the SAME background — paired, so the error bar is the sd of per-background differences rather than sd/√n.
- Intuition
- Necessity and sufficiency can come apart, and a neutral background isolates the second: whatever the motif does there, it does without help. The backgrounds preserve dinucleotides because yeast promoters are heavily biased toward poly(dA:dT) — a mononucleotide shuffle would destroy that too, turning a composition control into a composition change.
- Cost
- 11,600 forward passes, minutes.
- What would refute it
- CACGTG is its own reverse complement, so the forward and reverse arms must score identically and the script raises if they do not. Poly(dA:dT) has no scramble at all and is reported as degenerate rather than as a measurement.
What the answer tells us Sufficiency in a neutral background and necessity in context are different properties, and a motif can have either without the other. What this arm alone establishes is what a motif carries when nothing else is present.
Each consensus from the paper's dictionary implanted at the centre of 200 dinucleotide-shuffled backgrounds, paired against the same background without it (Global Importance Analysis, Koo & Ploenzke 2021). The tick on each bar is that motif's own scramble — the same bases in a shuffled order, so same composition and no motif. A bar that does not clear its own tick is measuring composition.
Two controls make this a measurement rather than a ranking. CACGTG is its own reverse complement, so implanting it forward and backward must score identically — the generator raises if the two arms differ, and it is the only check that catches a mis-wired implantation path, because a wrong one still produces entirely plausible numbers. And poly(dA:dT) has no scramble: every permutation of a homopolymer is itself, so its composition control does not exist and it is reported as degenerate rather than as a measured zero.
The model recovers known regulatory logic from sequence alone. Among the transcription-factor motifs the three strongest activators are Cbf1, Phd1, Tye7p — and all three are E-box motifs containing CACGTG. The three strongest repressors are Ume6.2, PAC motif (Dot6), Dot6p: Ume6's URS1 site, and Dot6 twice over, since the PAC motif is what Dot6 binds — together the ribosome-biogenesis repressor module. Nothing told the model any of that.
The two splice signals are held out of that ranking, and they split. A branch point implanted alone is the most repressing element in the whole set (-0.13277, z -13.834) while a 5′ splice site raises predicted coverage (0.03504, z 4.209). Neither is a regulatory motif and the quantity being predicted is RNA-seq coverage, so they are reported separately rather than ranked against Rap1 and Ume6 — but the split is real, and it is the direction a spliceable intron would push each of them.
Necessary and sufficient are different claims
- Question
- Are the motifs the knockout sweep calls necessary the same ones implantation calls sufficient?
- Why this matters
- Necessity and sufficiency are routinely collapsed into one word, 'important', and the motifs where they part are the informative ones.
- Method
- The two-by-two of in-context necessity against in-background sufficiency, from the two arms measured above.
- Intuition
- Crossing the two arms puts every motif in one of four cells. A general regulatory factor should come out necessary and NOT sufficient — it works through the promoter around it — which makes that expectation testable rather than assertable.
- Cost
- Free. Both arms are already run.
- What would refute it
- The answer is largely no, and that IS the finding: Rap1 and Reb1 are necessary and not sufficient at z ≈ 0.2, and Ume6 is the mirror.
What the answer tells us The off-diagonal cells are the informative ones. A factor that is necessary and not sufficient is working through the promoter around it — which is what a general regulatory factor should look like, and what a single importance score would have hidden.
Crossing the knockout sweep in act 5 against the implantation above. Horizontal is necessity — what the prediction loses when a curated site of that factor is shuffled in its own promoter, over every swept site. Vertical is sufficiency — its margin over its own scramble in an empty background.
Reb1 and Rap1 are necessary and not sufficient, and they are the two most-swept factors on the page — 14 and 23 curated sites, among the largest knockout effects anywhere in act 5 — yet implanted into a neutral background neither is distinguishable from its own scramble. Both are yeast's general regulatory factors, and the model has them working through the promoter around them rather than alone. Ume6 is the mirror image: strongly sufficient, barely necessary in these 23 windows, which are mostly constitutive genes where a meiotic repressor has little to do.
Where does a motif work?
- Question
- Does a motif work anywhere, or only in a particular place?
- Why this matters
- A motif's effect is usually quoted as a single number, which presumes position does not matter. For a promoter element that presumption is precisely what is in question.
- Method
- Implant at every position across the window and subtract a SCRAMBLE implanted at the same position.
- Intuition
- Overwriting real sequence destroys whatever was there, and in a promoter that is exactly where the real sites are — so the artefact would look like the signal. Subtracting a scramble implanted at the same position cancels it.
- Cost
- 47,104 forward passes, 21 minutes.
- What would refute it
- Overwriting 8 bp of a real promoter destroys what was there — in a promoter, exactly where the real sites are. Without the scramble at every position that artefact would look like the signal.
What the answer tells us A motif whose effect depends on where it sits is a positional element and not merely a sequence one; a flat curve says the opposite. Both are answers, and the same-position scramble subtraction is what makes either trustworthy.
Panel 1 asked whether a motif does anything at all, dropped into empty DNA. This asks the question a promoter actually poses: the same motif walked across a real window 64 bp at a time, scored on that window's own gene, and plotted against distance to the transcription start site.
The motifs that were sufficient are the motifs with a position. Cbf1 reaches +0.09398 log₂ in the 500 bp before the TSS and -0.00188 inside the gene body — it works upstream and nowhere else, which is the same asymmetry the attribution metaprofile found by a completely different route (1.40× a gene's mean base upstream against 0.94× inside). Ume6 represses from both sides, more strongly inside. And Rap1 and TATA box — the two that panel 2 showed are necessary at their own promoters but not sufficient in empty DNA — are flat everywhere here too. Three methods, one answer.
And every curve goes flat beyond about ±2 kb, which is not something this measurement was set up to show. It is where act 2's receptive-field sweep independently put the median window's effective context — ±2,048 bp of the ±8,192 available. Implanting a strong activator outside that radius moves the prediction by nothing, because outside it there is nothing left to move.
Why a scramble is implanted at every position too
Overwriting eight bases of a real promoter destroys whatever was there. A bare implantation curve is therefore partly a map of what was DESTROYED rather than of what was added — and in a promoter that is exactly where the real sites are, so the artefact would look like the signal. At each position the same bases in a shuffled order are implanted as well, and the plotted quantity is the difference. Both arms overwrite the same span with the same composition, so the damage cancels.
Aligned on the direction of transcription, never on coordinates. txStart on the plus strand, txEnd on the minus, minus-strand profiles reversed. Without that flip the average puts promoters against terminators and flattens into something indistinguishable from a real null — a mistake this site made once already, with the attribution metaprofile.
Does the arrangement matter?
- Question
- Does the arrangement of two motifs matter, and is there helical phasing?
- Why this matters
- Grammar claims — that two sites cooperate at a helical spacing — are among the most attractive and least tested in regulatory genomics, and a periodogram will always return periods.
- Method
- Walk a motif pair 4–200 bp apart in four orientations, re-measuring the moving motif's SOLO effect at every separation so its positional dependence is not folded into the interaction.
- Intuition
- Helical phasing predicts something specific: cooperation recurring near every 10.5 bp, the turn of the double helix, because that is when two sites face the same way. A periodogram always returns periods, so the discriminating test is the explicit in-phase against anti-phase contrast — and any period near a simple fraction of the scan window is the ruler rather than the DNA.
- Cost
- 25,448 forward passes.
- What would refute it
- The explicit in-phase against anti-phase contrast is the primary reading, because a 61-point periodogram cannot separate 10.5 bp from its own sixth harmonic (n/6 = 10.17 bp). Measured ratio 1.094 — no helical grammar.
What the answer tells us A helical grammar would show as a periodicity surviving the solo-effect subtraction. Its absence, with short-range unphased interaction present, says this model has learned proximity rather than phasing.
One motif anchored at the centre, a second walked away from it one base at a time in all four orientations. Plotted is the interaction — the pair's effect minus each motif's own effect at its own position — so a flat line means the model simply adds the two up. Grey rules mark multiples of 10.5 bp, one turn of the double helix.
There is no helical grammar, and the constructive test agrees with the Hessian below. Across the six pairs the in-phase / anti-phase ratio has a median of 1.0938 and runs from 0.7423 to 1.3676 — both sides of 1.0, which is what noise looks like rather than a suppressed signal. Orientation barely matters either. What the sweep does find is short-range and unphased: the strongest interaction anywhere is Cbf1 × Cbf1 at 6 bp (0.20858 log₂), and Reb1 reaches TATA best at 31 bp.
Why the singles are re-measured at every separation, and how phasing is read
The second motif's solo effect is measured at each separation, not once. A motif's effect varies with position for reasons that have nothing to do with the other motif, and subtracting a single number would fold that positional dependence into the "interaction" — manufacturing exactly the structure this panel is looking for. Every configuration also runs on the same backgrounds, so the curve is paired across separations and background variation cannot appear as spacing structure.
Phasing is read two ways, because a periodogram alone has misled this page before. The explicit contrast compares separations near integer multiples of 10.5 bp (the two motifs on the same face of the helix) against those near half-integer multiples (opposite faces). The periodogram is reported with the harmonics of the scan window printed beside it — the Hessian run's apparent periods of 49, 73.5 and 36.8 bp turned out to be 147/3, 147/2 and 147/4, artefacts of the analysis window rather than anything about DNA.
All eighteen top periods here are scan-window harmonics — n/2, n/3, n/4, n/5, n/7 and n/9 of a 61-point scan — and 10.5 bp appears nowhere. But the limit of that reading has to travel with it: n/6 is 10.17 bp, within a third of a base of one helical turn, so a scan this length cannot separate helical phasing from its own sixth harmonic. That is why the explicit phase contrast is the primary reading and the spectrum is the secondary one.
This is the constructive counterpart to the Hessian panel below, which asked the same question as a second derivative at real sequence and answered no. A local curvature and a built arrangement are genuinely different questions; where they disagree, both numbers are reported rather than reconciled.
Which bases interact with a binding site
- Question
- Which bases interact non-linearly with a binding site?
- Why this matters
- Everything else on this page is first order, scoring each base on its own. A first-order method cannot express a pair of bases that only matter together.
- Method
- A Hessian-vector product, H·v = ∇⟨∇f, v⟩ — one extra backward pass per site, mapping every base that interacts with that motif.
- Intuition
- The second derivative at the reference is the cheapest object that can carry an interaction — one extra backward pass. Its symmetry is the only available check that the double-backward is wired correctly, since a mis-wired one has the right shape and magnitude.
- Cost
- 425 ms a site — 431 sites across 23 windows, 18.7 a locus on average.
- What would refute it
- H is symmetric, so ⟨H·v_A, v_B⟩ must equal ⟨H·v_B, v_A⟩; the worst residual over 431 sites is 4.2e-04. Nothing else catches a mis-wired double-backward — it has the right shape and magnitude and is simply the wrong quantity.
What the answer tells us A base that interacts with a site is one no first-order map can score correctly, which is the gap this panel exists to fill. Whether these particular second-order values predict real double edits is a separate question, and act 9 measures it.
Regulation is not additive — factors cooperate, compete, and care about spacing — and a first-order method cannot see any of that by construction. For a motif spanning [a, b) with indicator v, H·v = ∇⟨∇f, v⟩ is one vector over all 16,384 bases whose entry i says how much that motif's own sensitivity changes when base i changes. One extra backward pass — 425 ms — rather than the 16,384² Hessian.
A second derivative is not a double substitution, and act 9 now measures the gap. Every pair of positions in a locked panel was mutated together — 39,330 exact double mutants — and the residual left after subtracting both single effects is what this panel is trying to predict. The Hessian does not call it: median r = -0.0843 across 23 loci, against 0.3227 for separation alone. Read this panel as the local curvature map it computes, not as a prediction about what two edits do together. The calibration is in act 9.
The helical-periodicity prediction is refuted, and that is the result. If the model had learned B-DNA's 10.5 bp pitch, |H·v| would repeat with it. Pooled over all 431 ChIP-supported sites in the 23 windows it does not: dividing out the decay and comparing distances in phase with 10.5 bp against anti-phase gives a ratio of 0.948358 — if anything slightly lower in phase. A falsifiable prediction, tested, and answered no.
Why the periodogram's answer is the wrong answer
The spectrum of the detrended profile puts its strongest periods at 49 bp, 73.5 bp, 36.75 bp — which are 147/3, 147/2 and 147/4, harmonics of the 4–150 bp analysis window. Reporting them as periods would be reporting the ruler. The phase test above needs no spectrum and is not vulnerable to it.
The check that a mis-wired second derivative fails is symmetry. H is symmetric, so ⟨H·vA, vB⟩ must equal ⟨H·vB, vA⟩ for two sites in one window. The worst residual over every pair is 4.2e-4, which is float32 arithmetic. Nothing else catches this: a wrongly wired second derivative has the right shape and the right magnitude and is simply the wrong quantity.
What would the model build?
- Question
- What would the model build, given a promoter to improve?
- Why this matters
- Reading what a model finds important is a weaker test than asking it to construct. A model can attribute sensibly and still hold no usable notion of what makes a promoter strong.
- Method
- Greedy ascent on the prediction, one substitution at a time, with the motifs it invents matched against the database afterwards.
- Intuition
- Greedy ascent verified by real forward passes turns the question into a construction. The control decides what it means: the same ascent on shuffled DNA says whether the motifs it builds are about yeast or about the search.
- Cost
- 5.5 minutes — 12 greedy substitutions a window chosen from 20 candidates, and the same ascent again on shuffled DNA as the control.
- What would refute it
- The same ascent on dinucleotide-shuffled DNA builds 54 motifs in 23 of 23 windows against the real arm's 18 in 11 — the control wins 3.0× raw, and the panel says so. What survives is that the achievable gain is set by headroom, r = −0.873 against starting expression.
What the answer tells us What the model builds is a statement about what it takes to make a promoter strong. The control decides whether that statement is about yeast: if shuffled DNA yields as many motifs, the ascent is finding the structure of the search rather than of the genome.
The generative question, and the only one here that lets the model choose the sequence. Starting from each real window, the 12 single-base edits that most raise its own gene's predicted expression — proposed by the gradient, then verified by real forward passes, because a gradient is a local slope on a saturating function and this page has already measured where that fails.
What it finds is headroom, not promoter quality. The gain runs against starting expression at r = -0.8726: the five quietest genes gain 2.7958 log₂ from twelve bases while the five loudest gain 0.3906. HOP2 and GAL3, both repressed in this baseline, move most; TDH3 and PDC1, already near the top of the model's range, barely move at all. A promoter the model thinks is maximal cannot be improved by editing it.
And the interesting claim does not survive its control. It is tempting to report that the ascent invents recognisable regulatory motifs — it does, 18 of them across 11 windows. But running the identical ascent on dinucleotide-shuffled DNA of the same composition builds 54 in 23 of 23. Raw, the control wins 3×; normalised by how much expression each arm actually gained, the two are within 25% (0.6913 against 0.5498 motifs per log₂). Either way, building a Reb1 site is a property of the ascent and the base composition — not evidence about promoters.
Why the control is the whole experiment
Without it this panel would have published a wrong result. "Given a free hand, the model rebuilds real yeast regulatory elements" is the sentence the raw numbers invite, and it is exactly the kind of claim that reads as a finding while being a fact about the search. The shuffled arm costs the same compute and settles it.
The control is not perfect and the page says so: shuffled DNA starts near silent, so it has far more room to gain and passes through more sequence space, which is why the raw counts are confounded and the normalised comparison is the one quoted. Neither version supports the strong claim.
Edits are not restricted to the promoter. Where they land relative to the TSS is recorded rather than assumed, so "the model edits promoters" would be a result rather than a constraint built into the experiment.
What does the model think a motif is? discovered, not matched
- Question
- Left to itself, does the model's mutagenesis contain repeated sequence patterns — and are they the motifs a database would name?
- Why this matters
- Every motif on this page so far was supplied — from a database or from the paper's figure. Whether the model's own attributions contain repeated patterns is a different and harder question.
- Method
- TF-MoDISco in the small (Shrikumar et al. 2018): pull high-|saliency| windows out of the mutagenesis planes at each pre-registered width (11 and 15 bp), cluster them on their mean-centred contribution blocks over both strands, build each cluster's PWM from the bases its seqlets contain, and only THEN match against JASPAR.
- Intuition
- Any clustering returns clusters and a lower threshold returns more, so the threshold has to come from the control before either arm is clustered — and the control has to perturb the model's input, not only the projection.
- Cost
- Free. No model and no GPU — the raw mutagenesis planes are already on disk.
- What would refute it
- The identical pipeline on dinucleotide-shuffled sequence, and the shuffled arm sets the clustering threshold. It found almost as much: the control is close enough that the headline has to be stated carefully, and it is.
What the answer tells us Repeated patterns in the model's own attributions would be motifs it discovered rather than motifs it was handed. Whether they clear a null that recomputes the model on shuffled input is settled below in this same panel, and until they do the clusters describe the clustering rather than the sequence.
Every other panel on this page starts from an annotation and asks whether the model agrees with it. That is the right way round for validation and the wrong way round for discovery — it can only ever find motifs already in a database, and it cannot say what the model thinks a motif is.
| pattern | seqlets | windows | information | closest JASPAR matrix |
|---|---|---|---|---|
AAATAAAAATA | 20 | 7 | 4.95 bits | SMP1 (MA2706.1) · r = 0.87 |
TATATAAAATT | 11 | 8 | 7.00 bits | HAP1 (MA0312.3) · r = 0.97 |
AAAAATATATA | 10 | 7 | 4.32 bits | MOT2 (MA0379.1) · r = 0.91 |
Both seqlet widths were declared before the run, and both are here. At 11 bp the real planes give 3 patterns and the shuffled control 2; at 15 bp the real planes give 3 patterns and the shuffled control 3. At the wider width the control ties exactly — three against three, all matched. Running a second width after a weak first result is only defensible because the grid was fixed in advance and neither cell is dropped; no third width was added afterwards.
Read the control before the table. Of 1,047 seqlets pulled from the real planes, 41 fall into 3 groups large enough to be called a pattern — the rest are singletons or pairs below the size threshold, which is what a strict similarity cut is for. All 3 of the patterns match a JASPAR matrix. The same pipeline on dinucleotide-shuffled sequence yields 2 patterns, and 2 of those match too. Three against two is not a result. What does separate them is sharpness — the best real pattern carries 7.00 bits against the control's 3.91 — but every pattern in both arms is AT-rich, and a dinucleotide shuffle preserves AT-richness by construction.
So the honest statement is narrow, and the second width narrows it further. At a threshold set by the control arm's own 99.9th percentile, what the model's highest-attribution 11-mers have in common is base composition rather than motif syntax — and at 15 bp, where the control matches the real arm pattern for pattern, there is no separation left to interpret at all. That is a real finding about where a fixed-width, single-pass clustering of 1,000 seqlets can and cannot reach — not evidence that the model has no motif vocabulary, which the knockout sweep and the sufficiency panel both show it does.
Why the control sets the threshold, and the PWM bug that came first
The threshold is not a choice. Any clustering returns clusters, and a lower correlation returns more, so picking one by eye and then reporting the count would be choosing the answer. The threshold used — 0.679 — is the 99.9th percentile of the shuffled arm's own pairwise similarity distribution. A cluster is therefore, by definition, a set of seqlets more alike than the control arm essentially ever manages — which is a statement about that arm, not about chance. An earlier version of this sentence said "essentially impossible"; the arm is finite, and until the matched null below it was also not a model-conditioned null at all.
The control shuffled the letters, not the model's input. It reused the mutagenesis planes computed on the real window and shuffled only the sequence used for saliency projection and reference-base choice — so the model never ran on the null sequence, and the two arms were not estimating the same quantity. The matched null fixes that: contribution maps recomputed on dinucleotide-shuffled input, the same quantity by the same route on both sides, with each shuffle draw clustered separately so the two arms face the same pool size — clustering is superlinear in seqlet count, and pooling the draws would have handed the null a five-fold advantage that dividing by the number of draws does not undo. It gives 0 clusters on real sequence (1,037 seqlets) against 3 ± 1 per draw over 5 draws of about 1,114 seqlets each. The real arm is NOT ahead — a sharper negative than the projection-only control gave.
Where that difference lives matters, and it is not in the bulk. The two arms' own pairwise-similarity distributions are indistinguishable — median 0.1051 on real sequence against 0.1047 on shuffled input — so the real arm's seqlets do not resemble each other less on the whole. The entire difference is in the tail above the threshold, over 5 draws. That is enough to say the real arm does not exceed a matched null, which is the claim the panel needed and did not have. It is not enough to say shuffled input is genuinely richer, and the page does not say so. Note also that this arm uses gradient × input, the contribution type both sides can be recomputed on; the ISM-based grid above still has no valid null, so no cell of this experiment has a model-conditioned null supporting a real-arm excess.
The first version's PWMs were 0.00 bits. They were built as a softmax of the mean-centred contributions, which is near-uniform whenever the contributions are small — and they usually are. The "consensus" was the argmax of noise, and it still matched JASPAR at r = 0.93, because Pearson normalises away amplitude. The PWM is now counted from the bases the seqlets actually contain, in the orientation the clustering chose, which is what makes information content mean anything.
What a seqlet is. A 4 × 11 mean-centred block — the hypothetical contribution of every base at every position — not a per-position score. Clustering on the score alone discards which base was doing the work, so AAATTT and TTTAAA would cluster together.
8Does it agree with biology outside the model?
Everything up to here is a statement about a frozen model. A model can be internally coherent, survive every control on this page, and still have learned something that has no counterpart in a cell. These three ask whether it lines up with evidence it was never given: phylogeny in the species channel, condition-dependence over an induction time course, and purifying selection in segregating yeast variation.
Tell the model it is a different fungus
- Question
- Does the species channel carry anything, with the DNA held byte for byte?
- Why this matters
- The model takes a 165-way species one-hot the paper never explains. Whether it encodes anything, or is a nuisance input the network learned to ignore, changes how every other result here should be read.
- Method
- Sweep all 165 species one-hots on a fixed S. cerevisiae promoter and rank the prediction.
- Intuition
- Holding the DNA byte for byte and changing only the label isolates the channel. The control that decides whether it means anything is between-locus rank correlation — a per-species bias would produce the same ordering everywhere.
- Cost
- 165 forward passes a locus, 1.6 min for the set.
- What would refute it
- The between-locus rank correlation is the control: median 0.33, so the orderings genuinely differ and this is not a per-species bias. The tidy story that the exceptions are the repressed genes is NOT supported — rank against log baseline coverage is only ρ = −0.34.
What the answer tells us A channel that reorders predictions carries information; one that only scales them is a gain. The between-locus control separates the two, and only the first would license reading the channel as phylogeny.
Shorkie's input carries a 165-dimensional one-hot naming which Saccharomycetales genome a sequence came from — an architecture no other sequence-to-function model has, and a counterfactual available here and nowhere else. The DNA below never changes; only the declared species does.
Told the truth about what it is reading, the model predicts the most expression. S. cerevisiae ranks first of 165 at 19 of 23 windows — probability 6.37e-39 under a null where the ordering is unrelated to the sequence — and sits above the median species at 21 of 23, by a median of 0.912 log₂ units. The exceptions are GAL1, GAL3, HOP2, GLK1, and all of them top out on Yarrowia lipolytica, the most distant lineage in the set.
The tidy explanation of those exceptions does not survive checking. Three of the four — GAL1, GAL3 and HOP2 — are the genes repressed in this baseline, which invites the story that the species channel only helps where a gene is being transcribed. Across all 23 windows that story is not supported: rank against log baseline coverage is only ρ = −0.34, a weak tendency rather than a rule, and the fourth exception breaks it outright. POP4 is among the quietest windows in the set and ranks first; GLK1 is far louder and ranks 162nd of 165. The pattern is suggestive and it is reported as suggestive.
The control that decides whether any of this means anything
A per-species bias would look identical at one locus. If every window ranked the 165 the same way, this would be measuring which channels simply predict more — nothing about the sequence. The between-window rank correlation is a median 0.3323 (range -0.6928 to 0.9834), so the orderings genuinely differ: the model reads the same promoter differently depending on the species it is told to be.
A second control comes free. 8 of the 165 rows are Yarrowia lipolytica strains — separate channels the model was never told to treat alike. Their spread is 9.5% of the full 165-species range, which bounds this method's noise from inside the data.
The index is confirmed, not assumed. The paper's own species list has exactly 165 rows and row 109 is Saccharomyces cerevisiae — matching the index this site had previously established by peak magnitude alone, against shuffled and random-sequence controls.
Do the drivers change as the response unfolds?
- Question
- Do the sequence drivers change as an induction response unfolds?
- Why this matters
- The model predicts a time course, and it is natural to read that axis as mechanism. Whether the sequence drivers actually move along it has to be measured rather than assumed.
- Method
- Attribution at early and late timepoints, per REGULATOR rather than on the timepoint means.
- Intuition
- Timepoint means average over hundreds of regulators and are near-identical by construction, so a per-regulator contrast is the only level at which a real change could show.
- Cost
- Free. Precomputed per locus.
- What would refute it
- On the timepoint means early-against-late is r ≥ 0.9995 with 99% of the top 500 bases shared — no effect at all. Per regulator it is real, MSN2 at GAL1 reading 0.215 with 50.2% shared.
What the answer tells us A driver that moves between timepoints is a sequence element whose relevance is condition-dependent. On the timepoint means the question is unanswerable because those means are near-identical, so the per-regulator level is where any real dynamics has to appear.
The 3,053 induction tracks resolve into 335 regulators over 13 timepoints — a sparse grid, 3,053 of 4,355 possible combinations, so a timepoint does not carry every regulator. Taking one regulator's own tracks early and late and attributing each over this gene body asks whether the model reads a different part of the sequence as the response develops.
It does, and only where the pairing is real. Across 26 regulators × 23 windows the median early-to-late correlation is 0.9973 — but the floor is 0.2149, at MSN2 on GAL1, where only half the strongest 500 bases survive. MSN2 and MSN4 are the paralogous general-stress factors that bind the same STRE motif, and they rank one and two independently; GAL1 and GLK1, where they move most, are both glucose-repressed.
Why this is per regulator and not per timepoint
The obvious version fails, measured twice. The 13 timepoint mean coverage tracks correlate at ≥ 0.9923 with each other, and their gradients are worse still — early against late attribution comes out at r ≥ 0.9995 with 99% of the top 500 bases shared. T0 alone carries 384 tracks while the later timepoints carry 64–77 each, so those means are over far fewer regulators, and averaging 300 induction experiments leaves the shared baseline transcriptome. The variation is real and lives in the individual tracks, where a single gene body spans 43.7× across them.
The replicate floor is on count, not on the timepoint label: T15, T30, T45, T60 and T90 carry 64–77 tracks against T0's 384, and a thin mean makes a noisy gradient. At two or more replicates 26 regulators qualify; at three, only eight.
Does selection agree with the model?
- Question
- Does natural selection avoid the bases the model says matter?
- Why this matters
- Every result above is about the model. Whether any of it corresponds to selection in real yeast populations needs evidence from outside it.
- Method
- Paired at the base: the observed allele against the two alternates nature did not choose at the SAME position, so position, context, gene and expression are all held fixed and no null model is needed.
- Intuition
- If the model has learned what constrains a sequence, the bases where it predicts a large effect should be bases where surviving variation is milder, because selection has already removed the costly alternatives. The prediction is directional — the observed allele is the gentler one — so a sign test is the right instrument, and the skew in the per-site ratios is why a mean is the wrong one.
- Cost
- Free. No model run — the 2,319 UCSC variants that fall inside the windows mutagenesis covers, scored against the shipped planes.
- What would refute it
- A sign test, because the ratios are skewed: for missense the ratio of means is 0.99 while the median per-site ratio is 0.90. The reference base matched sacCer3 at 2,319 of 2,319, which is what makes the coordinates trustworthy.
What the answer tells us This is the only evidence on the page from outside the model. It is an alignment with selection, not a demonstration that the model has captured the mechanism selection acts on.
2,319 real segregating variants from UCSC's evaSnp8 fall inside the windows mutagenesis covers — and those packs hold all three substitutions at every base. So for each variant the allele that actually segregates can be compared with the two alternates at the same base that nature did not choose. Paired, so position, context, gene and expression level are held fixed by construction and no null model is needed.
They agree, modestly and consistently. The observed allele is the milder one 55.8% of the time for missense (sign test z = 2.54), 55.4% for synonymous (z = 3.3) and 54.6% for non-coding (z = 2.74), with a median per-site ratio near 0.90 in all three. The model has never been shown a variant; this is purifying selection on expression, visible through a network trained only to predict assays.
Why a sign test, and the control that makes the coordinates trustworthy
A sign test, because the ratios are heavily skewed. For missense the ratio of mean effects is 0.99 while the median per-site ratio is 0.90 — a few sites where the observed allele is much the louder pull the mean up. The design is paired, so under neutrality the observed allele is the milder one exactly half the time, and a distribution-free test on that is the one the data already earns.
The reference base named in each variant record matched sacCer3 at 2,319 of 2,319 checked. A one-base coordinate convention error, or an allele reported against the other strand, would otherwise have produced a clean and completely meaningless answer.
The reference row of a mutagenesis plane is zero by construction, so it is excluded rather than averaged into the unobserved pair — including it would drag every comparison toward zero and manufacture the result.
9How much of this should you believe?
Every number above comes from one training fold, one baseline and a method that had never been scored against an exact edit. These four panels measure that, and they are the only ones on the page whose subject is the page itself. 1 shipped attribution method fails its own benchmark, and the default integrated-gradients baseline is not the best of five — both were found here rather than argued.
Does the signal survive retraining? 85% of strong bases
- Question
- Would a different training run of the same architecture call the same bases important?
- Why this matters
- Every headline on this page comes from one checkpoint. If a different training run of the same architecture ranks different bases, the page is reporting a property of one optimisation rather than of the architecture or the data.
- Method
- Every exact single-base effect in a locked panel per locus — the 128 strongest by mutagenesis plus 128 matched on distance to TSS — recomputed under all 8 released folds and on both strands. Fold and strand are kept as separate axes and never pooled.
- Intuition
- Eight released folds are eight draws from one training procedure, so their disagreement estimates training variance directly. The panel must be selected on one fold and scored on the others, or the selection is being reported back to itself.
- Cost
- 17,664 substitutions × 8 checkpoints × 2 strands, about an hour on an M1 Pro, plus 385 MB of checkpoints that are public and were simply never downloaded.
- What would refute it
- The panel is chosen from f0, so f0's own agreement with it is a tautology: f0 selects and is reported separately, and every cross-fold statistic below is computed over f1, f2, f3, f4, f5, f6, f7 only. The matched controls are what "how many survive" is measured against.
What the answer tells us A base that keeps its sign across eight independent training runs is a property of the architecture and the data; one that does not is a property of a single optimisation. The matched controls set the scale — 54% survive by construction, so 85% is what selecting on mutagenesis actually bought. The pairwise overlap adds the caveat that the headline cannot: direction is stable, ranking much less so.
The eight folds are eight different models, and this was measurable all along. Both audit documents record training-fold uncertainty as unestimated. It is not unestimable — all eight checkpoints sit at a public URL, and on one locus they predict g across a range of a factor of 1.68 in coverage. What follows is what happens to the page's own top-ranked bases when the training run changes.
All eight checkpoints, one row each
The summary above is a median over these. Each row is one released training run and each column one locus, coloured by how far that fold's own prediction sits from the across-fold mean for that locus. Deviation rather than the raw score because the loci span 5.8 to 14.9 log₂ units — on an absolute scale every column would be one flat colour and the differences between checkpoints, which are the whole subject, would be invisible.
How much do two checkpoints agree?
Sign is stable; ranking is much less so, and the two are easy to conflate. The headline above says 85% of the strongest bases keep their direction across folds. But asking which bases each checkpoint puts in its top 1% is a different question, and the answer is 41–55%. Two checkpoints trained on the same data agree about roughly half of what matters most. A single-fold ranking is therefore a weaker object than a single-fold sign.
Strand and fold disagreement are comparable in size, and neither is noise. The median spread across folds is 0.0085 against a median strand deviation of 0.0080, a ratio of 0.943×. The model was trained with augment_rc: false, so forward and reverse are genuinely different states and averaging them is a test-time choice, not a symmetry being exploited. Pooling the two axes into one error bar would hide both.
The three gates, fixed before the run
Sign agreement > 0.8, measured over the evidence folds only; the page's median is 1.000. Top-1% overlap > 0.4 between every pair of evidence folds on the whole-window gradient — 164 of 16,384 bases — where the median is 0.485. And a 95% interval across folds excluding zero. A base is called fold-stable only if it clears the first and the third; the second is a per-locus statistic.
Why the panel is selected from f0 and f0 is then excluded. The bases worth asking about are the ones the site's claims rest on, and those come from f0's plane. Measuring f0's agreement with a panel f0 chose would report the selection back to itself. So the selection and the evidence are deliberately different checkpoints.
Which attribution methods survive an exact edit? 3 of 4
- Question
- Do the page's attribution maps rank bases the way exact mutagenesis does — and do they beat a baseline that knows nothing about the model?
- Why this matters
- The site shipped five attribution surfaces and had never scored one against an exact edit. Without that, choosing between them is an aesthetic judgement.
- Method
- Rank agreement with the exact ISM planes, each method scored at its OWN native resolution with the ground truth pooled to match, plus a deletion curve: rank the bases, substitute the top-k to their worst alternative, and measure the real fall in g.
- Intuition
- Exhaustive mutagenesis is the finite-difference contrast every other method approximates, so it is the natural reference. Rank agreement alone is not enough: a map can correlate well and still order the wrong bases first, which only an intervention reveals.
- Cost
- Minutes. The exhaustive mutagenesis ground truth for all twenty-three windows was already on disk; what was missing was the interventions that score against it.
- What would refute it
- A method that cannot beat GC content, distance to TSS or a random ranking is demoted on this page rather than defended. Attention rollout is included knowing it is unsigned token mixing at 128 bp and has no reason to pass.
What the answer tells us A method that clears the baselines has earned the right to be read as an attribution; one that does not is a picture. This is the verdict act 6's copy now defers to, and it is why attention rollout is described there as token mixing rather than importance — it finds 4.0% of the achievable damage against 15.3% for a random ordering.
Scoring a 64 bp method per base measures its resolution, not its faithfulness. Occlusion and rollout resolve 64 and 128 bp, so the mutagenesis truth is pooled to each method's own grid before the ranks are compared. The deletion curve gives them no such allowance — within a block it can only order bases by position, which is exactly the handicap their resolution imposes on a reader.
| method | resolution | rank rs vs exact ISM | deletion AUC | verdict |
|---|---|---|---|---|
| exact mutagenesis — the ceiling | per base | 0.415 | 1.000 | the scale |
| integrated gradients | per base | 0.469 | 0.897 | promote |
| gradient × input | per base | 0.696 | 0.858 | promote |
| occlusion | 64 bp | 0.623 | 0.286 | promote |
| conv-stem response — baseline | per base | 0.004 | 0.166 | — |
| random ranking — baseline | per base | 0.002 | 0.153 | — |
| distance to TSS — baseline | per base | 0.555 | 0.117 | — |
| GC content — baseline | per base | -0.037 | 0.059 | — |
| attention rollout | 128 bp | 0.529 | 0.040 | demote |
The two columns measure different things, and the ceiling row shows it. The oracle is ordered by how far g falls when each base is substituted to its worst alternative — the quantity the deletion curve applies — so it bounds that column at 1.000. Its rank agreement is only 0.415, below gradient × input, because the ground truth for that column is the largest effect a base can have in either direction. The two orderings part company wherever a base's biggest effect is an increase. An earlier version ranked the oracle by |effect| and integrated gradients scored 1.0053 against it — a ceiling a method can exceed is a mislabelled competitor, not a ceiling.
1.000 is ranking by the answer. Every curve is divided by what the exact mutagenesis ranking itself achieves, so the column reads as the fraction of achievable damage a method found. The bar is the strongest baseline, at 0.166 — GC content is a real predictor of where a promoter is, and distance to TSS still better, so a method that merely finds promoters clears a random ranking without demonstrating anything about the model. The conv stem is filed with them deliberately: it is linear and unnormalised, so its filters are a basis any invertible recombination leaves unchanged.
attention rollout does not clear the baselines. attention rollout finds 4.0% of the achievable damage, against 15.3% for a random ranking. Attention rollout is unsigned token mixing at 128 bp — an architectural quantity describing what the transformer can read, not an attribution — so this is the expected result rather than a surprise, and the panel in act 6 says so on its face. It is reported here because a method nobody scores is a method nobody can demote.
Every fold is scored against fold f0's exact mutagenesis, not its own. That is the only exhaustive mutagenesis that exists — per-fold planes would be about 52 GPU-hours — so the fold arm asks a specific question: do the page's f0-derived conclusions transfer to another checkpoint? The page's claims are f0-derived, so that is the right question, but it is not "each checkpoint against its own ground truth", and one consequence shows in the numbers: the oracle is a strict ceiling only for f0, and another fold's ranking can edge past it.
The verdicts hold across the eight training folds, which is what makes them verdicts. The whole benchmark was re-run under every released checkpoint: integrated gradients, gradient × input, occlusion clear the baselines in 8 of 8; attention rollout in none. No method is split across folds, so no verdict here rests on which checkpoint was used. The spread across checkpoints differs by method: gradient × input ranges 0.742–0.914 while attention rollout ranges 0.026–0.040, so how much a method's score depends on the training run is itself a property of the method.
Does the map read the model, or the input?
An attribution that survives having the network destroyed is not reading the network. This is the Adebayo sanity check: randomize the weights from the head backwards, stage by stage — permuting each tensor's own values, so every parameter keeps its exact distribution and only the arrangement is lost — and watch the map decorrelate from its intact self. Gradient × input falls from 0.96 with only the head destroyed to 0.25 with the whole network destroyed, over 6 loci.
| randomized through | branch | |rs| vs intact |
|---|---|---|
| head | decoder | 0.962 |
| decoder1 | decoder | 0.605 |
| attn_out8 | bottleneck (skips bypass it) | 0.511 |
| attn_out1 | bottleneck (skips bypass it) | 0.342 |
| block7 | encoder, feeds a skip | 0.271 |
| stem | encoder, below every skip | 0.249 |
The branch column is why one number would have misled. Destroying the entire transformer takes the map only from 0.60 to 0.34 — not to zero — because the three decoder skips are fed by block5, block6 and block7 and carry real signal around the bottleneck. A pooled score would have read that architectural fact as the check failing. It does not settle at zero either: gradient × input on a fully randomized network still shares 0.25 with the intact map, which is what the one-hot input geometry contributes before any learning.
What a deletion curve is, and why completeness is quoted twice
The deletion curve is the only part of this that is an intervention. Rank agreement is a correlation between two maps. The curve takes the method's own ranking, substitutes the top k bases to the alternative the exact plane says is worst, re-runs the model forward and reverse, and records what actually happened to g. A map can correlate well and still rank the wrong bases first.
Integrated gradients' completeness gap is 0.0166 absolute and 0.8% relative, and both are quoted because the relative figure is unreadable on its own: where the target gap is near zero a miss of 0.04 reads as several hundred percent, which is a fact about the denominator.
Is the IG baseline doing the storytelling? best: dinucleotide shuffle
- Question
- The site ships one integrated-gradients reference — all four DNA channels zeroed. Is that choice carrying the result?
- Why this matters
- A path method integrates from a reference, and that reference is a modelling choice usually made for convenience. If the choice carries the result, the map describes the baseline rather than the sequence.
- Method
- The same 32-step path integral under five legal reference families on the same scalar and the same loci, every one preserving the species one-hot and none assuming channel 4 is a mask. Scored by agreement with exact mutagenesis and by real deletion passes.
- Intuition
- Every family is scored on the same scalar and the same loci, so none gets an easier target. A reference the sequence model finds impossible is one the expression model was never trained near, which makes plausibility measurable instead of arguable.
- Cost
- 23 loci × 5 families × 2 strands × 32 steps, about ten minutes.
- What would refute it
- Shorkie_LM's own masked negative log-likelihood scores each reference: a baseline the sequence model finds impossible is one the expression model was never trained near. And strand-wise maps are compared BEFORE averaging — a family that only looks stable once the two arms are averaged has failed, not passed.
What the answer tells us A path attribution inherits its reference, so this decides how the integrated-gradients lane may be read. The shipped default does not win and has the weakest strand agreement in the table, so integrated gradients ships here labelled reference-sensitive rather than as the map.
| reference | rank rs vs ISM | deletion AUC | completeness | strand agreement | LM surprise (bits) |
|---|---|---|---|---|---|
| dinucleotide shuffle | 0.188 | 0.955 | 0.0182 | 0.738 | 1.96 |
| expected gradients | 0.397 | 0.952 | 0.0242 | 0.449 | 1.96 |
| mononucleotide shuffle | 0.189 | 0.934 | 0.0235 | 0.738 | 1.98 |
| all-zero DNA — shipped | 0.469 | 0.899 | 0.0166 | 0.358 | — |
| another real window | 0.205 | 0.829 | 0.0343 | 0.748 | 1.76 |
The shipped baseline does not win — dinucleotide shuffle does. All-zero DNA reaches 0.899 of the achievable damage against 0.955, and its forward and reverse maps agree at only r = 0.36 — the weakest in the table. A family that needs the two arms averaged before it looks stable has failed the transparency requirement, not passed it. Integrated gradients therefore ships labelled reference-sensitive, with the menu above exposed, rather than quietly defaulting to one path.
Across all 8 checkpoints, one thing is robust and one is not. The shipped all-zero default wins in 0 of 8 folds — every shuffle-based family beats it under every checkpoint, so "the default is not the best" is a property of the reference rather than of one training run. Which family is best is not robust: the winner moves between dinucleotide shuffle, mononucleotide shuffle, expected gradients across folds (dinuc, mono, expected, mono, expected, mono, mono, expected), and their medians sit within 0.031 of each other. So this panel recommends moving OFF the all-zero baseline, and does not recommend a specific replacement. The single-fold result named a different winner from the eight-fold median, which is the clearest argument on this page for why the fold axis was worth the compute.
Rank agreement and deletion damage part company here, and the shipped baseline is where they part hardest. All-zero DNA has the best rank correlation with exact mutagenesis in the table — 0.469 — and nearly the worst deletion AUC. Correlating with the ground truth and ordering the bases whose mutation actually hurts are different achievements, which is the same dissociation attention rollout shows one panel up. The most in-distribution reference does worst of all: a real window shares so much with the sequence being explained that the path between them carries little.
Real yeast sequence costs Shorkie_LM 1.7648 bits a base under its own masked schedule, which is the number every reference in the table should be read against. An all-zero window is not merely unusual sequence — zeroing the four DNA channels is literally how this model family masks a position, so that reference asks the expression model to integrate a path from a fully masked window.
Does the second derivative predict a real double edit? no — r = -0.0843
- Question
- The epistasis panel in act 7 is a Hessian — curvature at the reference. Does it predict what two actual substitutions do together?
- Why this matters
- Act 7's interaction panel is a second derivative. Whether curvature at a point predicts what two discrete edits do together is an assumption this page had never tested.
- Method
- Every double substitution over 20 positions a locus — nine per pair — evaluated exactly and rc-averaged, with the inclusion–exclusion residual as the measured interaction. Singles are recomputed in the same run so both terms of the subtraction share a convention.
- Intuition
- A Hessian is the limit of an infinitesimal perturbation; a substitution is a finite jump to another vertex of the simplex. Recomputing the singles in the same run is what makes the subtraction meaningful, since both terms must share one convention.
- Cost
- 39,330 exact double mutants across 23 loci, about half an hour.
- What would refute it
- Additive predicts exactly zero and is the null the Hessian has to beat. Separation alone is scored too, in case proximity is all the second derivative is tracking. The recomputed singles are checked against the shipped mutagenesis planes; the worst drift is a free correctness check on the whole path.
What the answer tells us Whether curvature at the reference predicts real double edits decides how act 7's interaction panel may be read. At r = -0.0843 it does not, so that panel is a local curvature map. The interactions are real — a median residual of 0.051× a single-base effect — but which pairs interact has to be measured, not inferred from a second derivative.
A second derivative is not a double substitution. The Hessian is curvature at the reference one-hot; a real double edit is a finite jump to another vertex of the simplex, and nothing guarantees the first predicts the second. Half the positions are the strongest bases and half are matched on distance to TSS, so "how much interaction is there" has something to be compared against.
The Hessian does not call the measured interaction, and two trivial predictors do better. Over 39,330 exact double substitutions the Hessian's median correlation with the measured residual is -0.0843, ranging -0.3223 to 0.6826 and positive at only 8 of 23 loci. Separation alone reaches 0.3227 and is positive at every locus, and the summed magnitude of the two single effects does better still. The pipeline is not what failed: the singles recomputed for this benchmark reproduce the shipped mutagenesis planes to 0, so both terms of the subtraction share a convention exactly.
Re-measured under every checkpoint, the failure is not a property of one training run. The Hessian's correlation with the measured residual is positive in 0 of 8 folds (-0.11, -0.17, -0.17, -0.03, -0.06, -0.08, -0.02, -0.00), and separation alone beats it in 8 of 8. A predictor that a trivial baseline outperforms under most checkpoints is not a predictor of this quantity. The fold arm uses a smaller panel than the headline — 10 positions a locus against 20, so 9,315 pairs a fold rather than 1,710 — because the fold question is whether the calibration's sign survives retraining, which a smaller panel answers at a quarter of the cost. The headline above is the full panel on f0.
A curvature is not an edit, and that is the whole finding. The Hessian is the second derivative at the reference one-hot; a double substitution is a finite jump to two other vertices of the simplex. Interactions here are real but small — a median residual of 0.051× a median single-base effect — and the second-order term at the reference does not point at them. The act 7 panel should be read as a local curvature map, which is what it computes, and not as a prediction about combined edits. An earlier single-locus probe of this benchmark gave r = +0.79 on 90 pairs; the full grid is 39,330, and a correlation from one locus is not the correlation.
This is a calibration, not a spacing scan. The separations here are wherever the strong bases happen to be, so they are not uniform and no periodicity claim can be read off them. The constructive spacing sweep in act 7 is where that question belongs — and it already reports no helical grammar.
—How do I read any of this?
Every panel above collapses a multi-dimensional tensor to something drawable, and they do not all collapse the same axis the same way — a difference that is invisible in the drawings and changes what a comparison between two of them means. This appendix states each collapse, the layer shapes, and where the checkpoint and the paper disagree.
Every dimension this page collapses
Nothing here is drawable at its real size. A stage is [channels × positions] with up to 16,384 positions; the output is [896 × 5,215]. Every panel therefore collapses at least one axis, and — this is the part worth knowing — they do not all collapse the same axis the same way. Position is max-pooled for display, because the question there is “did any neuron fire here”, and sum-pooled for relevance, because relevance is additive and a mean would make a coarse stage look quiet for being coarse. Both are right; the difference is invisible unless it is written down.
| where | tensor | collapsed how |
|---|---|---|
| Flow canvas, layer raster | [C × T] per stage | max-pool position → 128 |
| Conv-stem profile | [96 × 16,384] | max-pool → 1,024, plus a per-filter max |
| Attention map | [8 × 4 heads × 128 × 128] | mean over heads |
| Output-head raster | [896 × 5,215] | mean over the tracks in each of 4 assay groups |
| Every-output-track heatmap | 5,215 rows → pixels | max over the tracks a pixel row covers |
| Channel margin (trace) | |∂f/∂a ⊙ a| [C × T] | sum over position |
| Position margin (trace) | |∂f/∂a ⊙ a| [C × T] | sum over channel, then sum-pool → 128 |
| Layer relevance map | the two margins | outer product, normalised to sum 1 — the one estimate here |
| Stage-stack row | position margin [128] | scaled by its 99th percentile, then centred on its own mean |
| Neuron traces | [C × T] | none — top 8 channels drawn whole |
| Attention rollout | 8 × [128 × 128] | ∏(½A + ½I) row-normalised, then mean over the region's rows |
| Occlusion track | [256 × 896] | mean over the traced output bins |
| Mutagenesis logo | [4 × W] | mean-centre across bases, then project — keeps 1 of 4 |
| Coverage curve | 5,215 tracks | mean over the group's tracks (3,053 for RNA-seq) |
| Attribution target | 5,215 tracks × 896 bins | mean over 384 T0, sum over gene bins, then log2 |
| Annotation mask (enrichment) | features → [16,384] | union, not a count — overlapping features do not stack |
| Annotation mask (neuron classes) | [16,384] → [128] | mean — the covered fraction of each cell, never a max |
| Enrichment ratio | signal × mask | mean |·| inside ÷ mean |·| overall; null by circular shift |
| Knockout sweep | k shuffles per site | mean ± sd over shuffles, never a single draw |
| Method logo stack | method → letters | none for per-base methods; a band at its own step for the rest |
The last five rows are new, and the second is the one most easily got wrong: a 7 bp binding site covers 5% of a 128 bp cell, so pooling the annotation by max would mark that whole cell as annotated and make every class look identical — producing numbers that mean nothing while looking exactly like numbers that do.
Two rows are worth reading against each other. The layer relevance map is the only place on this page where a number is reconstructed rather than measured, and the neuron traces exist directly beneath it as the un-collapsed answer to the same question. Where they disagree, trust the traces.
Which model this is, and how close the score is to the paper's
One fold, and that is the paper's choice. Figure 4's ISM reads .../f0c0/part<N>/scores.h5 — a single model, fold 0 of the cross-validation, which is the checkpoint this page runs. The repository does ship an 8-fold ensemble loader, but it belongs to Figure 7's eQTL analysis, not to these logos. Averaging eight folds here would move this page away from the figure it reproduces, so it does not.
The score is a different formula from the paper's, and they agree. Figure 4 averages per-track logSED across the 384 _T0_ tracks; this page computes logSED of the track-averaged coverage. Those are not the same quantity — a mean of logarithms is not the logarithm of a mean — so it was measured rather than assumed. Over 240 real substitutions in TDH3's promoter the two agree to r = 0.999987, with a largest difference of 0.00033, the same strongest substitution, and a scale ratio of 1.005. That is far below what the shipped uint8 packs can even represent. The reason is that the T0 tracks are strongly correlated; for a heterogeneous track set the two would genuinely diverge.
What is not checked. The paper's own ISM arrays live on the authors' cluster and are not in the public repository, so nothing here is a numeric comparison against published values. The claim is that the recipe matches — the fold, the track subset, the gene-body slice, the rc-averaging, the mean-centring and the one-hot projection — and that the drawing reproduces the published figure.
Every layer, with its tensor shapes
Paper vs checkpoint
8 places where the published Methods and the released f0 checkpoint disagree. This page follows the checkpoint, because that is the model actually running.
| Input channels | paper: the helper library says 4 DNA + 166 species (ensemble.py:17-20) checkpoint: 4 DNA + 1 unused + 165 species. The corpus table settles it: 1/80/165/1361 species give num_features 6/85/170/1366 — always n+5, and six loaders inject exactly that. |
|---|---|
| The fifth channel | paper: named a "special channel" once, and never written checkpoint: identically zero at every position in every shipped inference path; the LM masks by zeroing the four DNA channels instead |
| Attention heads | paper: 8 heads checkpoint: 4 heads (r_w_bias is [1, 4, 1, 64]) |
| Residual block | paper: BatchNorm → GELU → Conv1D(5 bp) checkpoint: adds a second, pointwise Conv1D(1 bp) and a learned per-channel Scale |
| Decoder stage | paper: BatchNorm → GELU → Dense → UpSampling → skip merge checkpoint: ends each stage with a SeparableConv1D(3 bp) |
| Filter progression | paper: 96 → 384 in 32-filter steps checkpoint: 96, 128, 160, 192, 256, 320, 384 |
| Parameter count | paper: 13.7 M checkpoint: 14,253,567 |
| Output track order | paper: RNA-seq (3,053), 1,000-strain (1,014), ChIP-exo (1,128), ChIP-MNase (20) checkpoint: ChIP-exo 0–1127, ChIP-MNase 1128–1147, RNA-seq 1148–4200, 1,000-strain 4201–5214 |