Released models¶
Three models are public. All live in the public bucket gs://seqnn-share
under the shorkie_models/ prefix, are downloadable over plain HTTPS, and are
catalogued with sizes and MD5 checksums in data/manifest.json.
Model |
Folds |
Training |
What it’s for |
|---|---|---|---|
Shorkie_LM |
1 |
Masked LM ( |
Masked-base prediction, embeddings, and the |
Shorkie |
8 |
Fine-tuned from Shorkie_LM ( |
The main model. Coverage prediction and variant effects |
Shorkie_Random_Init |
8 |
Same data/architecture from random init ( |
The ablation isolating the contribution of LM pretraining |
Download¶
data/download.sh --models all # all three (~1.4 GB)
data/download.sh --models lm # Shorkie_LM only
data/download.sh --models finetuned # the 8-fold Shorkie
data/download.sh --models random_init # the ablation
data/download.sh --minimal # just the 8 Shorkie folds (~0.46 GB)
Every file is MD5-verified against the manifest as it downloads.
Direct links¶
On-disk layout¶
data/download.sh writes the layout the loaders expect:
<release_root>/models/
├── shorkie_lm/
│ ├── params.json
│ └── train/model_best.h5
├── shorkie_finetuned/
│ ├── params.json
│ ├── targets.txt # the 5215-track sheet
│ └── train/f{0..7}c0/train/model_best.h5
└── shorkie_random_init/
├── params.json
└── train/f{0..7}c0/train/model_best.h5
Note
params.json sits at the model-dir root, while the checkpoint is under
train/. Some author work-dirs put params.json under train/ too, so the
examples accept either location.
Architecture¶
All three share one unet_small_bert_drop architecture — the ablation deliberately
holds it fixed:
Input:
(16384, 170)— channels 0–3 are DNA one-hot, channels 4–169 are species identity (column 114 = S. cerevisiae).Output:
(1, 1, 896, 5215)— 896 bins of 16 bp, one channel per track.~13.7 M parameters.
The committed training configs under scripts/02_train/ match the
released params.json field-for-field; a test in the repository pins that
equality so the published recipe cannot drift from the published weights.
Choosing a model¶
Use Shorkie unless you have a specific reason not to. Shorkie_Random_Init exists to answer “how much did pretraining help?” — it is not a better or newer model, and because it never saw the LM it is generally weaker. Shorkie_LM is the right choice only when you want representations or masked-base probabilities rather than expression predictions.