The Live Variant Playground

Shorkie · Chao et al. 2025 · fold f0, 14,253,567 parametersLanguage model →LabPost

1What does it predict, and which bases drive it?

A 16,384 bp window in, coverage over 896 bins and 5,215 tracks out. Everything below shares one horizontal axis: pick a locus and a region in the bar above, then read straight down — prediction, attribution, the four methods, the curated evidence for what is really there, and finally the highlighted window again at single-base resolution.

Sequence

Each locus is a 16,384 bp window — the model's full context — centred on the named gene. Six of them are the windows Figure 4 of the paper prints, at the published coordinates; the rest are classic high-expressers and the galactose regulon.

Predicted coverage — the full model

showbinding sites

What is actually there, from the databases. Genes, CDS, introns, uORFs, tRNA and snoRNA genes, replication origins, LTRs and transposons come from SGD; regulatory regions from ORegAnno; binding sites from the Harbison 2004 / MacIsaac 2006 conserved regulatory code and from a JASPAR PWM scan. Provenance is drawn, not just recorded: a ChIP-supported site is solid, a conserved-only call is hollow, a PWM match is a hairline — three different strengths of evidence that must not be read as one. A grey cap marks a feature clipped by the window edge, so that edge is never mistaken for a real boundary.

zoom · the highlighted window above

The whole window, as the paper draws it. Red dashed boxes are matched motifs and splice landmarks; click one to zoom the logo below onto it. The navy track is the gene model in the paper's own IGV style.

Drawn exactly as the paper draws it — the real DejaVu Sans Bold glyphs at the paper's own per-letter offsets, scaled 1.35× on both axes, A green / C blue / G orange / T red, stacked largest-magnitude nearest the axis with positives up and negatives mirrored below. Letters overflow their column and touch; that density is the published look, not a bug. The colours are fixed in every theme, because a figure's base colours are part of the figure.

Two sources, and they are not the same method. Mutagenesis is the paper's Figure 4 quantity: every substitution actually run, then mean-centred across the four bases and projected on the reference, so the height reads as “how much the base that is there matters” rather than “what this one substitution did”. It now covers every base of the window — all 49,152 substitutions per locus, on both strands — so it is the primary track here and needs no caveat about where it stops. Gradient × input is a local linear sensitivity — it is identically zero at the three bases not present, because the input is one-hot, so it draws one letter per position. The paper's published figures use no gradients at all; its single gradient routine is unreached scaffolding that still declares a 4-channel input the 170-channel model cannot accept. Gradient × input is here because it is instant, not because it is what the paper did.

What the whole window shows that a promoter window cannot. Running all 16,384 bp on both strands for every locus makes one result visible that no single window does: the sign of the strongest substitution is predicted by the gene's expression state in 19 of 19 windows. The three genes repressed in this baseline — HOP2 (meiosis-specific), GAL3 and GAL1 (galactose regulon, in glucose) — are the only three whose largest substitution raises predicted expression (+1.38, +1.12, +0.78 logSED, with 20–33 of their top 40 positive). The other eleven are all negative. For a gene already being transcribed the model's biggest lever is breaking what drives it; for a silent one it is breaking what holds it down.

Why that is not just an artefact of the log scale

The objection. logSED is a log ratio, and those three genes have by far the lowest baseline coverage — 77, 213 and 502 against up to 48,547. A given absolute change in predicted coverage moves logSED 161× more at GAL3 than at TDH3, so low-baseline genes should show large log effects whatever the biology.

The control, from the data itself. Baseline does not separate the two groups cleanly, and where it overlaps the sign still tracks regulation: DTD1 (baseline 299) and MMS2 (556) sit between GAL3 (213) and GAL1 (502), and both are firmly negative (−0.45, −0.42). Two pairs at matched baselines with opposite signs — so the sign is about regulation, not scale.

The magnitude is a different matter, and the orderings invert. As a fraction of its own baseline the largest lever on the page is HOP2's +160%. In absolute predicted coverage HOP2's best substitution is worth +125 while PDC1's costs −8,841, seventy times more. Neither reading is wrong; they answer different questions. logSED is deliberately a ratio — that is what makes a silent promoter and a maximal one comparable — so "a stronger effect" is meaningless here without saying stronger relative to what.

What the old 500 bp window would have missed. It covers 3.1% of the sequence and holds a median 17.6% of the window's total |logSED| — genuinely well chosen, 5.7× enriched. But the strongest substitution falls outside it in 4 of the 14 windows, including HOP2's +1.38, the largest effect in the whole set.

Which model this is, and how close the score is to the paper's

One fold, and that is the paper's choice. Figure 4's ISM reads .../f0c0/part<N>/scores.h5 — a single model, fold 0 of the cross-validation, which is the checkpoint this page runs. The repository does ship an 8-fold ensemble loader, but it belongs to Figure 7's eQTL analysis, not to these logos. Averaging eight folds here would move this page away from the figure it reproduces, so it does not.

The score is a different formula from the paper's, and they agree. Figure 4 averages per-track logSED across the 384 _T0_ tracks; this page computes logSED of the track-averaged coverage. Those are not the same quantity — a mean of logarithms is not the logarithm of a mean — so it was measured rather than assumed. Over 240 real substitutions in TDH3's promoter the two agree to r = 0.999987, with a largest difference of 0.00033, the same strongest substitution, and a scale ratio of 1.005. That is far below what the shipped uint8 packs can even represent. The reason is that the T0 tracks are strongly correlated; for a heterogeneous track set the two would genuinely diverge.

What is not checked. The paper's own ISM arrays live on the authors' cluster and are not in the public repository, so nothing here is a numeric comparison against published values. The claim is that the recipe matches — the fold, the track subset, the gene-body slice, the rc-averaging, the mean-centring and the one-hot projection — and that the drawing reproduces the published figure.

How a motif box is placed, and why some are hollow

A solid box is a motif; a hollow box is a database region. Curated binding-site calls come from a PWM scan under a conservation filter, not from matching a consensus string — so most of them do not contain their factor's canonical motif. Measured across these windows, only 22.8% of 670 curated calls do. Crucially their coordinates are not wrong: where the consensus is present it sits at offset 0 of the call. So a box is drawn on the match where one exists, tight to the motif and with the matched bases printed beneath it — the reference-DB convention of the paper's own Figure 4F — and on the database's region otherwise, hollow and labelled as such. Two different claims, drawn differently.

Splice and codon landmarks are spans, not a fixed window. They used to be boxed as six bases centred on the junction, which framed AAGGTA where DTD1's donor motif is GTATGT. Each landmark now carries the span of the motif it names — donor 6 bases, acceptor 2, codons 3, branch point 7 — mirrored on the minus strand, and a test decodes the sequence inside every drawn box to confirm it.

The consensuses come from Figure 4H, the paper's motif dictionary. Recording them explicitly caught one already: this page had Fhl1 as GTAAACA, which is not the paper's Fhl1 motif, and was boxing a position 27 bp from where Figure S19B puts it.

How the logo is drawn, and why only one letter appears per position

The geometry is the paper's, not an approximation of it. Real DejaVu Sans Bold outlines for A/C/G/T with the paper's own hand-tuned per-letter x-offsets (A −0.350, C −0.366, G −0.384, T −0.305 — not half-advances), scaled by globscale = 1.35 on both axes, which is what makes a stack summing to S fill height S and what makes adjacent letters touch. Colours are the paper's saturated X11 values #008000 / #0000FF / #FFA500 / #FF0000, fixed in every theme because a figure's base colours are part of the figure. Letters are stacked descending by magnitude, positives up from zero and negatives mirrored below, with y-limits at the data min and max each padded by 5% of max|v|.

Mutagenesis is the paper's Figure 4 quantity, in three steps (fig4_common.py:227-230, and identically in two other files): average logSED over the T0 tracks → mean-centre across the four basesproject on the reference one-hot. Step two is what turns "what this substitution did" into "how much the base that is there matters"; step three is why exactly one letter survives per position. Given the shipped plane P with the reference cell zero, that value is exactly −(Σ P)/4.

Gradient × input also draws one letter per position, but for a different reason: the input is one-hot, so gradient × input is identically zero at the three bases that are not there. That is a property of the method, not a rendering choice.

Occlusion resolves 64 bp, so every base inside a window carries that window's value and the logo reads as blocks. Its sign is flipped to match the others: occlusion measures what is lost, so a base that matters has a negative logSED, and the logo convention is that up means "raises the prediction".

drag across the method tracks to zoom

The same window, as letters, for every method that has them. A logo of the whole 16,384 bp window is not a drawing that exists — one base is 0.078 px — so the strip above stays a signal and the letters live in a zoom. Drag across any method track to move it. Only gradient × input and integrated gradients are per-base, so only those two become letters; attention rollout resolves 128 bp and occlusion 64 bp, and they are drawn as bands at their real resolution rather than stretched into letters they cannot support.

Four methods, one axis, one annotation. They do not measure the same thing and are not expected to agree everywhere: mutagenesis substitutes one base and re-runs the model; gradient × input is a local linear sensitivity at the reference base, mean-centred across the four bases as Borzoi does; integrated gradients accumulates that gradient along a path from an all-zero baseline and is the only one here whose values sum to the prediction difference — a property it would lose if it were mean-centred, so it is not; occlusion ablates a whole 64 bp stretch and measures what the model loses. Where they agree the result is robust; where they diverge, the divergence is the finding. A method that covers only part of the window is drawn in place with the rest left blank, which is what the paper does for its own partial windows.

Traceback. Gradient × input: how much the prediction over the selected bins moves per unit of each input base, times the base that is actually there. It is a local sensitivity, not a decomposition — the values do not sum to the prediction. Dragging is exact rather than interpolated, because gradients superpose: the attribution for a set of bins is the sum of their individual gradients. Dragging resolves to 128 bp; picking a gene gives single-base resolution.

How the four method tracks are computed

All four answer the same scalar, so they can be laid on one axis: f = log2( Σgene bins mean384 T0 tracks coverage + 1 ). That is the paper's own logSED target (ensemble.py:97-104): mean over tracks, sum over bins, +1 pseudocount, log base 2. Being a log ratio is what makes a silent promoter and a maximal one comparable.

  1. Gradient × input. One backward pass of f w.r.t. the one-hot input gives [16384 × 4]. Mean-centre across the four bases (g − meanACGT g) — the Borzoi convention — then multiply by the input and sum the base axis. Because the input is one-hot, three of every four values are exactly zero, so this is one number per position. Cost: 1 backward pass. Reads as: if I nudged this base up a little, how would the prediction move. Fails when: the model is locally flat but globally sensitive — a saturated promoter can show near-zero gradient at a base that matters enormously.
  2. Integrated gradients. The same gradient, accumulated over 32 steps along a straight path from an all-zero-DNA baseline (species channel kept, because a baseline with no species is not a sequence this model was trained to see), then multiplied by x − baseline. Not mean-centred, deliberately: completeness — that the values sum to f(x) − f(baseline) — is a telescoping integral of the raw gradient, and centring destroys it. Measured, centring pushed the error from ~0.05 to 8–650%. Cost: 32 forward+backward per region. Reads as: this base's share of the whole difference between the real sequence and an empty one. Unique property: it is the only method here you can check against its own total, and the panel prints that check.
  3. Attention rollout. The eight attention matrices, each mixed half-and-half with the identity to account for the residual stream, renormalised and composed: ∏ (½A + ½I). Row i then reads as where position i's representation came from. Cost: free — the maps already ship. Reads as: what the transformer can see for this region. Not an attribution: it says what could influence what, never what changed the prediction, and it is unsigned by construction.
  4. Occlusion. Zero the four DNA channels across a 64 bp window and re-run; the change in f per output bin is that window's effect. Cost: 256 real forward passes × 2 strands. Reads as: what the model loses without this stretch. Fails when: two stretches are redundant — remove either alone and nothing happens, which occlusion reports as "neither matters".

Both strands, averaged, for all four, matching Borzoi and the --rc every published Shorkie ISM run passes. Worth knowing what that is: this model was not trained with reverse-complement augmentation (augment_rc: false in all four params.json), so it is not rc-equivariant — measured on TDH3 the target reads 15.60 forward against 14.23 reversed and the two gradients correlate at 0.31. Averaging them is a test-time augmentation, not a free symmetry.

How much do they actually agree? Measured on TDH3's own gene body, at the 64 bp grid occlusion resolves: gradient × input against integrated gradients r = 0.874, gradient × input against occlusion r = 0.901, integrated gradients against occlusion r = 0.856. The three full-window methods agree strongly, which is the reassuring part. Against mutagenesis, per base over its own 500 bp, gradient × input falls to r = 0.658 — and that is expected rather than alarming: a gradient is a local linear sensitivity and a substitution is a finite jump to a different base. Where a promoter is saturated the two genuinely differ.

One thing the curve above does not share. The drawn coverage is the 3,053-track RNA-seq group mean; every attribution scores the 384 _T0_ subset the paper uses. Measured, the two correlate at r = 1.0000 and differ by 1% at the peak, so the distinction does not affect reading — but it is a distinction.

What full-window mutagenesis costs, and why it was once refused. 16,384 × 3 substitutions on both strands is 98,304 forward passes a locus and 1,376,256 for all fourteen. An earlier round priced that at 39.6 h and dropped it — but that was measured through onnxruntime on the CPU, on a graph whose batch axis is pinned at 1, so it baked in both the slowest engine available and the impossibility of batching. Neither limit belongs to the model. Re-measured on the same machine's GPU at batch 32, with the output head sliced to the 384 tracks actually scored: 10.47 ms a forward pass — and a substitution costs two of them, because every published Shorkie run averages both strands. Measured over the run that produced the packs on this page: 23.8 ms a substitution, 19.5 minutes a locus, 4.6 h for the fourteen loci that run covered. The numbers on this page are that run. The engine change is not a precision compromise either — against the CPU the GPU agrees to 6.6 × 10⁻⁷ relative, three orders of magnitude tighter than the fp16 graph the earlier packs were built from.

The curve is already here. Every preset locus was run offline at the full 16,384 bp context and the predictions ship with the page, so there is nothing to wait for. Loading the model — a 28.6 MB download and about 17 s of WebAssembly per run — is what gives you the live layer activations, sequence editing and motif knockouts. 896 bins of 16 bp across the window's 14,336 bp interior. Purple bars are annotated ORFs. A dashed line, once you mutate, is the unedited reference. Measured coverage is the second curve where it is loaded: Shorkie's own BigWigs binned exactly as its training labels were — 16 bp sums, soft-clipped, 1,024 bp cropped from each end — so the two curves are the same quantity and Pearson r means something. The bucket they live in is requester-pays, so this site ships none of them; scripts/shorkie/make_truth.py builds the overlay from a local copy.

Every output track

One row per predicted experiment, in the order the released targets sheet lists them, so the four assay blocks read as bands. Click a row to plot that single track below — the overview above averages thousands of experiments, which is a curve no instrument would produce. Where rows outnumber pixels each drawn row is the maximum over the tracks it covers.

What drives what

Every input window against every output bin. Each column is a 64 bp stretch of the input, ablated; each row is one of the 896 output bins; the colour is what that bin loses when that stretch is gone. Blue is a loss, red a gain. This is the only genuinely two-dimensional view on the page — every other method collapses to a single profile over the input — and it is measured, not inferred: 256 real forward passes a locus.

Read it as a map, and read both parts. The dark diagonal is local effect — a window damaging the output directly above it — and per cell it is by far the strongest thing here. But it is also narrow: one window of 256. Summed across the row, the local footprint accounts for only 0.2–5.6% of the most damaging window's total effect, so the diagonal is where the model is most intense and the off-diagonal is where most of what it does actually lives. Neither statement alone is the truth. Off the diagonal, a vertical stripe is an input stretch many outputs depend on and a horizontal stripe is an output bin that reads widely.

The edges are the margins. The strip along the top is the column sum — how much each input window matters summed over every output — and the strip down the right is the row sum, how much each output bin depends on the sequence at all. Click a row to read that one output bin's input profile, or a column to see which outputs that window drives; click again to clear. Bins differ in expression by orders of magnitude, so the checkbox rescales each row to its own range when the raw map is dominated by the loudest ones.

Ablation zeroes the four DNA channels, which is exactly how the paper's language model masks a position and is indistinguishable from a run of N. That is a different question from the motif knockouts above, which shuffle: a shuffle keeps base composition and asks whether the arrangement matters, while zeroing asks whether the stretch carries information at all.

How the occlusion map is computed

For each of 256 windows of 64 bp: zero the four DNA channels across that window, run the model, and record log2(alt + 1) − log2(ref + 1) for every one of the 896 output bins. One forward pass answers a whole row, which is why a complete [256 × 896] map costs 256 passes — 27 s a locus, 54 s with both strands — and why this is the cheapest exact method here.

Per cell the unit is logSED, a log2 fold change, so bins of wildly different expression are comparable within a row. The default colour scale is shared across the whole matrix for that reason; the checkbox rescales each output bin to its own range, which trades the between-bin comparison for the within-bin one and is worth it when a few loud bins dominate.

The margins are plain sums of |logSED|: the top strip is each input window summed over all 896 bins, the right strip each output bin summed over all 256 windows.

The diagonal is not the identity line. The two axes cover different spans — the input is the whole 16,384 bp window, the output only the cropped 1,024–15,360 interior — so it is a line of slope <1 that reaches neither corner. It is drawn dashed where it actually falls.

Zeroing is not shuffling. Zeroing the DNA channels is how the paper's language model masks a position and is indistinguishable from a run of N: it asks whether the stretch carries information at all. The motif knockouts above shuffle, which preserves base composition and asks whether the arrangement matters. Different questions, and they can disagree.

What occlusion cannot see: redundancy. If two stretches carry the same information, removing either alone changes nothing, and this map reports both as unimportant.

2What happens inside?

The same prediction, from the inside. Watch a sequence pass through all twenty stages, open any one of them, then trace one output region back through every layer to the bases.

The network, end to end — watch a sequence pass through it

showingSame layout either way: height is positions, width is channels.

Each block is one stage, sized by its real dimensions on the log scales named above. Click a stage to inspect it. Before you run the model the blocks are empty outlines; afterwards they fill with the activations from that same forward pass.

Switch to “relevance” and the same blocks show one region instead of the whole window. Trace a region below, and every stage repaints with what it contributed to that region alone — every other output bin masked out. Both of the map's margins are exact: the per-channel one and the per-position one are each a row-sum of precomputed groups, because gradients are linear in which outputs you select. Its interior is their outer product, which is the only reconstruction consistent with both margins if channel and position act independently — a real assumption, and the reason the panel names it.

This layer, in detail

    Trace a region — layer by layer, back to the sequence

    Five questions, in order. Each one narrows the last: which stages carry this region, then which neurons inside the strongest of them, where those neurons fire, what they respond to in the curated annotation, and finally what the transformer could have readregardless of what moved the prediction.

    1 Which stages carry this region?

    Longer bar, more of this region's relevance passes through that stage. Click a stage to open it in This layer, in detail above; click one of its channel chips to plot that neuron in step 2.

      Each bar is a mean over the stage's channels, never a sum: summing would rank a 1,536-channel stage above a 384-channel one on width alone, which is a fact about the architecture and not about this region. Stages whose activations live on their own tensors report own tensor rather than zero — "not measured here" is not the same claim as "contributes nothing".

      2 Which neurons, exactly?

      The top eight channels of the selected stage, each drawn as its real activation profile — nothing reconstructed. Enrichment is the share of a channel's activity inside the traced region over the share of the window that region occupies, so 1.0× means it fires there no more than anywhere else and the interesting rows are well above it.

      Each trace is scaled to its own range, so heights are not comparable between rows — the fires pair gives the real range. A faint rule marks zero, drawn only where the channel actually crosses it; several fire entirely negative. These traces are the un-collapsed answer to the same question the relevance map asks in This layer, in detail above: that map reconstructs a stage's [channels × positions] relevance as the outer product of two exact margins, so its margins are measured and its interior is an independence assumption. These are measured throughout. Where the two disagree, trust these.

      3 Where do those neurons fire?

      One row per stage, the window along x. Red where that stage draws more than its own average, blue less, neutral where it has no positional preference. Read it top to bottom: structure appears only from the transformer on.

      The early residual blocks come out close to neutral, and that is a real property rather than a flat drawing: pooled to 128 positions, a convolution over 16,384 bp is spatially almost uniform. Structure appears where the receptive field becomes the whole window and the model can choose where to look. Each row is the factorised combination — per-channel relevance times per-position activation — not a per-position gradient, because the shipped pack carries relevance summed over position. It answers “where do the channels that matter fire”, which is not quite “which positions the gradient flows through”.

      4 What are those neurons responding to?

      The same neurons as step 2, scored against the curated annotation instead of against the region: for each feature class, which of this stage's relevant channels fire on it more than chance. This is what turns “channel #14 matters here” into “channel #14 responds to binding sites”.

      Scored on the 32 channels with the highest exact relevance for this region, not all of them — the question is what the neurons that matter here respond to. Both the activation profile and the annotation are pooled to the packs' common 128 positions, so one cell is 128 bp: a channel cannot be shown to respond to a 7 bp site, only to the 128 bp neighbourhood containing it. That resolution limit is why a high ratio here is a lead, not a conclusion.

      How a neuron is matched to a feature class

      Same statistic as the enrichment table, applied one channel at a time. The channel's real activation profile [128] plays the role of the signal; the annotation mask, pooled from [16384] to [128] as the covered fraction of each cell, plays the role of the weight. The ratio is mean |activation| where the class is, over mean |activation| everywhere, and the null is the same deterministic circular shift.

      Pooling by fraction, not by max. A 7 bp site covers 5% of a 128 bp cell. Taking a max would mark that whole cell as annotated and make every class look identical once pooled, which is the failure that would make this panel meaningless while still producing numbers.

      What it is not. A channel that responds to a class is not a detector for it: the pooled resolution is 128 bp, the classes overlap each other, and nothing here controls for a channel simply firing where the gene is. Read a row as "worth looking at", then check the neuron's own trace in step 2.

      5 What could the transformer read at all?

      Attention rollout: the eight attention matrices composed, each mixed half-and-half with the identity to account for the residual stream. This is architecture, not attribution — it says what information could reach this region, never what changed the prediction, and it is unsigned by construction.

      It is a second, independent answer, and it needs no new data: the eight [128 × 128] maps already ship in every pack. Where it disagrees with gradient × input, the disagreement is the finding — a region the transformer can see but does not use is a different statement from one it cannot see.

      How the layer-by-layer trace is computed, and what each collapse throws away

      Every view in this panel starts from the same tensor and collapses it differently. For one stage, relevance = |∂f/∂a ⊙ a| over its activations a: [channels × positions] — how much each neuron, at each place, moved the prediction over the selected bins. That array is far too large to ship for every region, so what is precomputed is its two margins:

      1. Channel marginsum over position, giving [channels]. "Which neurons mattered, wherever they fired."
      2. Position marginsum over channel, then sum-pooled to a common 128, giving [128]. "Where in the window this stage drew from, whichever neuron did it."

      Both are exact, and both superpose: gradients are linear in which outputs you select, so an arbitrary contiguous region is the row-sum of the precomputed 112 groups of 8 bins — no model run, no interpolation. That is why dragging is instant and still exact.

      The map's interior is not exact. Reconstructing [channels × positions] from two margins requires an assumption, and the one used is independence: the outer product, normalised to sum to 1. It is the unique reconstruction consistent with both margins, and it is still an assumption. The neuron traces below the map exist precisely because they are not: each is one channel's real activation, straight from the shipped stage maps.

      The stage stack is the position margin, one row per stage, each row scaled by its own 99th percentile and then painted against its own mean — red above average, blue below, neutral where the stage has no positional preference. Per-row scaling because stages differ by orders of magnitude; against the mean because the early residual blocks are genuinely near-uniform once pooled to 128, and a from-zero ramp draws that truth as a saturated bar reading "maximally relevant everywhere".

      Three stages are missing from all of this, and that is not an oversight: the input one-hot, the conv stem and the output head live on their own tensors in the exported graph and have no per-layer relevance in the pack. They show their activations and say so.

      The enrichment number on each neuron trace is the share of that channel's Σ|activation| falling inside the traced region, divided by the share of the window the region occupies. 1.0× means the neuron fires there no more than anywhere else. Measured on block 4, the top-relevance channels come out 0.85–1.32× — so a channel's relevance to a region generally comes from what it computes there, not from firing only there.

      3Does any of it match real biology?

      The tests that can come out negative. Every curated binding site knocked out and measured, and the attribution scored against annotation the model never saw.

      What all nineteen windows say

      Every other panel in this act shows one window at a time. This one is the pattern across all nineteen, computed once from the same packs and the same null those panels use — so it cannot drift from them.

      Each row is one annotation class; the bar is the median across the windows that contain it, and the dots are the individual windows. A class present in two windows and a class present in fourteen are drawn the same width, so the count beside each row is what says how much to trust it.

      How the cross-locus summary and the TSS profile are computed

      Nothing here is new data. The per-class enrichment is the same statistic the panel below computes, run once per window offline and the medians taken; the knockout winners are read from the same -ko.json packs the sweep panel reads. It is a second view of numbers that already exist, not a second copy of them — verify_pipeline re-derives it from the packs and fails if the two disagree.

      The TSS profile aligns on the direction of transcription, not on coordinate order: the start is txStart on the plus strand and txEnd on the minus, and minus-strand profiles are reversed so "upstream" is the same side of the plot for every gene. Without that flip the average puts promoters against terminators and produces a flat curve that looks exactly like a real null result.

      Each gene is normalised to its own total before averaging, so a highly expressed gene does not dominate. The band is the spread across genes, not a confidence interval on the mean — these genes differ by orders of magnitude in expression and the spread is the honest statement of that.

      What it cannot tell you. 158 genes from nineteen windows is a small, non-random sample: they were chosen as classic high expressers, the galactose regulon and the paper's own figure panels. A metaprofile from them describes those genes, not the yeast genome.

      Every curated site, knocked out

      Not six hand-picked motifs — all of them. Every ChIP-supported binding site in this window was shuffled and the model re-run, five times per site, and what is reported is the mean logSED with its spread. One shuffle is one sample; a single draw presented as a measurement is exactly what this replaces. Scored over the window's own gene body on the paper's 384 T₀ tracks, both strands — the same quantity the interactive knockout reports, using the same seeded shuffle, so a swept value and a clicked one agree exactly rather than approximately.

      Where a gene has a well-known regulator, the sweep tends to find it, unprompted. Ranked by the size of the effect, the strongest site is RAP1 at the glycolytic genes TDH3 and PDC1 and TYE7 — also a glycolytic activator — at PGK1; FHL1 at the ribosomal protein gene RPL26A; and a GAL-regulon factor at both galactose genes, GAL80 at GAL1 and GAL4 at GAL3. At GAL1 the effect is unusually clean: the six largest sites in the window are all GAL4 or GAL80, the activator and the repressor of that regulon. Nothing in this pipeline knows which factor is supposed to matter — the sites come from a database and the ranking comes from the model.

      It is a tendency and not a law, and the table is ranked by magnitude so you can check it: at HOP2 and ACT1 the winner is STE12, which is a mating and filamentation factor rather than the regulator either gene is known for, and at KRE33 and DTD1 the winners are not characterised regulators of those genes at all. Sign is not reliable either — knocking out the GAL80 repressor site raises the prediction, which is the right direction for a repressor, but several activator sites also come out positive

      A site whose bar is inside its own spread did nothing measurable, and that is a result about that site rather than a broken button. Shuffling preserves base composition and destroys only the arrangement, so this asks whether the order of those bases matters — a different question from the occlusion map, which zeroes a stretch and asks whether it carries information at all. The two can disagree. A low-complexity site cannot be shuffled at all — one site across the fourteen windows is five Cs and an A, which has almost no distinct permutations — so those are labelled rather than reported as a null result. “Unshuffleable” and “the model ignores it” are different findings that produce the same number.

      How the sweep is computed, and why a mean over shuffles

      Per site: Fisher–Yates shuffle the bases inside the site, re-run the model, and record log2(Σ coverage + 1) over the gene body minus the same quantity on the unedited sequence. Repeat with k = 5 independent permutations and report mean ± sd. Cost is one forward pass per shuffle — 104 ms — so the whole ChIP-supported tier across all fourteen windows is a few thousand passes.

      Why not one shuffle. A single permutation is one draw from the distribution of "sequences with this composition". Two shuffles of the same site can differ by more than the effect being reported, so the single number the page used to show was a measurement with an unstated error bar of unknown size. The sd here is that error bar, measured.

      Why the gene body, not the peak. A 16 kb yeast window holds a dozen genes and the tallest is rarely the one whose promoter was edited. Measuring globally reports a number about an unrelated gene — on one window the global peak is 114.3 while the named gene's own body peaks at 7.8.

      What it cannot tell you. Shuffling a site does not remove the possibility that the model reads a different, overlapping element; sites here overlap frequently. And a null effect at a real binding site can mean the model is robust to it, that a redundant site elsewhere covers it, or that the site is not used in the condition the T₀ tracks represent.

      Does the attribution land on real biology?

      One number per method, per annotation class. For each class, the mean |attribution| on the bases that class covers, divided by the mean across the whole window — so 1.00× means no preference. The null comes from circularly shifting the annotation against the signal, which keeps the number of features, their lengths and their spacing exactly and destroys only their alignment; the shaded figure beside each ratio is that null's spread. Cells are shaded by how far the ratio sits from its own null, not by its size. The table measures every evidence tier whatever the lanes above are drawing — comparing the three tiers against each other is the point, and a drawing toggle should not silently change a statistic.

      Read the negatives too. A class at 1.0× is a real answer: it says the model puts no more weight there than anywhere else. And a method cannot resolve a feature finer than its own resolution — occlusion moves in 64 bp steps and attention rollout in 128 bp, so neither can say anything about a 7 bp binding site, which is why their columns are marked. This is a correlation against a stated null, not evidence the model uses the feature: a site that happens to sit inside a region the model attends to for unrelated reasons scores exactly the same.

      How the enrichment and its null are computed

      The statistic. Take the per-base signal s over the 16,384 bp window and a binary mask m for the class. The ratio is (Σ|s|·m / Σm) / (Σ|s| / L) — mean magnitude inside over mean magnitude everywhere. It is scale-free, so two methods with wildly different units are directly comparable, and it is a magnitude, so a class the model suppresses still reads as enriched. The mean signed value inside is reported alongside so that direction is not lost.

      The null is a circular shift, not a resample. The mask is rotated by each of 256 evenly spaced offsets and the ratio recomputed. Rotation preserves the feature count, every feature's length, and the spacing between them — so the null compares against an annotation that looks exactly like the real one but is not aligned to the signal. Resampling positions instead would compare against a feature set that does not resemble the real one at all, and would call almost everything significant. The offsets are deterministic (evenly spaced, zero excluded), so a published number is reproducible rather than a draw.

      The p is empirical, (#{null ≥ observed} + 1) / (256 + 1), so it is never exactly zero and its floor is 1/257 ≈ 0.0039. A cell at that floor means "no shift out of 256 reached this", not "p = 0".

      What it cannot tell you. Circular shifts are a weak null against broad positional trends: if a method's signal is concentrated near the window centre and a class happens to be too, the shift will not fully break that. It is also a per-class test with no multiple-testing correction across the table — read a single marginal cell with that in mind.

      4How do I read any of this?

      Every panel collapses a multi-dimensional tensor to something drawable, and they do not all collapse the same axis the same way.

      Every dimension this page collapses

      Nothing here is drawable at its real size. A stage is [channels × positions] with up to 16,384 positions; the output is [896 × 5,215]. Every panel therefore collapses at least one axis, and — this is the part worth knowing — they do not all collapse the same axis the same way. Position is max-pooled for display, because the question there is “did any neuron fire here”, and sum-pooled for relevance, because relevance is additive and a mean would make a coarse stage look quiet for being coarse. Both are right; the difference is invisible unless it is written down.

      wheretensorcollapsed how
      Flow canvas, layer raster[C × T] per stagemax-pool position → 128
      Conv-stem profile[96 × 16,384]max-pool → 1,024, plus a per-filter max
      Attention map[8 × 4 heads × 128 × 128]mean over heads
      Output-head raster[896 × 5,215]mean over the tracks in each of 4 assay groups
      Every-output-track heatmap5,215 rows → pixelsmax over the tracks a pixel row covers
      Channel margin (trace)|∂f/∂a ⊙ a| [C × T]sum over position
      Position margin (trace)|∂f/∂a ⊙ a| [C × T]sum over channel, then sum-pool → 128
      Layer relevance mapthe two marginsouter product, normalised to sum 1 — the one estimate here
      Stage-stack rowposition margin [128]scaled by its 99th percentile, then centred on its own mean
      Neuron traces[C × T]none — top 8 channels drawn whole
      Attention rollout8 × [128 × 128]∏(½A + ½I) row-normalised, then mean over the region's rows
      Occlusion track[256 × 896]mean over the traced output bins
      Mutagenesis logo[4 × W]mean-centre across bases, then project — keeps 1 of 4
      Coverage curve5,215 tracksmean over the group's tracks (3,053 for RNA-seq)
      Attribution target5,215 tracks × 896 binsmean over 384 T0, sum over gene bins, then log2
      Annotation mask (enrichment)features → [16,384]union, not a count — overlapping features do not stack
      Annotation mask (neuron classes)[16,384] → [128]mean — the covered fraction of each cell, never a max
      Enrichment ratiosignal × maskmean |·| inside ÷ mean |·| overall; null by circular shift
      Knockout sweepk shuffles per sitemean ± sd over shuffles, never a single draw
      Method logo stackmethod → lettersnone for per-base methods; a band at its own step for the rest

      The last five rows are new, and the second is the one most easily got wrong: a 7 bp binding site covers 5% of a 128 bp cell, so pooling the annotation by max would mark that whole cell as annotated and make every class look identical — producing numbers that mean nothing while looking exactly like numbers that do.

      Two rows are worth reading against each other. The layer relevance map is the only place on this page where a number is reconstructed rather than measured, and the neuron traces exist directly beneath it as the un-collapsed answer to the same question. Where they disagree, trust the traces.

      The network, end to end

        Paper vs checkpoint

        Seven places where the published Methods and the released f0 checkpoint disagree. This page follows the checkpoint, because that is the model actually running.

        Input channelspaper: the helper library says 4 DNA + 166 species (ensemble.py:17-20)
        checkpoint: 4 DNA + 1 unused + 165 species. The corpus table settles it: 1/80/165/1361 species give num_features 6/85/170/1366 — always n+5, and six loaders inject exactly that.
        The fifth channelpaper: named a "special channel" once, and never written
        checkpoint: identically zero at every position in every shipped inference path; the LM masks by zeroing the four DNA channels instead
        Attention headspaper: 8 heads
        checkpoint: 4 heads (r_w_bias is [1, 4, 1, 64])
        Residual blockpaper: BatchNorm → GELU → Conv1D(5 bp)
        checkpoint: adds a second, pointwise Conv1D(1 bp) and a learned per-channel Scale
        Decoder stagepaper: BatchNorm → GELU → Dense → UpSampling → skip merge
        checkpoint: ends each stage with a SeparableConv1D(3 bp)
        Filter progressionpaper: 96 → 384 in 32-filter steps
        checkpoint: 96, 128, 160, 192, 256, 320, 384
        Parameter countpaper: 13.7 M
        checkpoint: 14,253,567
        Output track orderpaper: RNA-seq (3,053), 1,000-strain (1,014), ChIP-exo (1,128), ChIP-MNase (20)
        checkpoint: ChIP-exo 0–1127, ChIP-MNase 1128–1147, RNA-seq 1148–4200, 1,000-strain 4201–5214