Changelog
v1.0.10
Incremental release (2026-07-30). Consolidates everything that accumulated
behind the v1.0.10 tag: the originally staged v1.0.10 work, a batch of
GitHub-issue fixes, a full-repository correctness audit, and a CDS-attribute
parity fix. Neither v1.0.10 nor the intermediate v1.0.10.1 was ever published,
so v1.0.9 is the baseline for every comparison below. As in v1.0.9, every
output-affecting default ships with an opt-out flag. See the project
CHANGELOG.md for the full list.
Changed (output-affecting defaults — results differ vs v1.0.9):
A third best-of-outcome merge candidate. Alongside {chained merge + ORF-rescue, Liftoff + ORF-rescue}, LiftOn now scores miniprot's native CDS-only model and adopts it only when its ORF-rescued protein identity is strictly higher than the two-way winner — so per-transcript identity never decreases — and never for an antisense hit. Pass
--no-miniprot-candidateto restore the two-way merge.A divergence-adaptive miniprot-only rescue floor. The rescue pass used a fixed protein-identity floor of 0.50; that floor is now lowered toward 0.30 as the DNA lift's gene recall drops, recovering more genuinely missing genes at large evolutionary distance while staying inert on same- and close-species lifts. Pass
--no-adaptive-rescue-floorto restore the fixed floor.Richer, spec-valid CDS rows. Rebuilt CDS lines lost the reference's descriptive attributes (
Dbxref,product,protein_id,gene,locus_tag, ...) and carried noIDat all. They now inherit those attributes and share oneID=cds-<transcript>— the spec's discontinuous-CDS form. Coordinates and the encoded protein are untouched; output grows 12–43%.LIFTON_NO_CDS_ATTR_CARRY=1andLIFTON_NO_CONTAINMENT_NORMALIZE=1reproduce the previous bytes.Coding transcripts are harmonized to ``mRNA``. The DNA-lift path preserved the reference featuretype — often the generic
transcript— while the miniprot path emittedmRNA, so one output labelled the same kind of feature two ways.LIFTON_NO_MRNA_HARMONIZE=1preserves the reference type.
Fixed (reported issues):
no such column: <id>from the DNA lift on strict SQLite builds (GH #35), which is why the same input worked on one machine and failed on another.A single unresolvable alignment killing an entire chromosome (GH #39).
Mixed-strand output from an antisense merge (GH #33).
UTR indels reported as frameshifts on close pairs (GH #46).
CDS emitted with no
ID(GH #32, GH #8) and coding transcripts typed inconsistently (GH #28).Header-only Liftoff output from an invalid or LFS-pointer minimap2 index (GH #57), and a
--streamingest crash on DuckDB 1.5.3/1.5.4 (GH #56).minimap2 is now preflight-checked at startup and documented as a requirement (GH #43, GH #11); the run summary is also written to
stats/summary.txt(GH #50).
Fixed (robustness):
A skipped locus no longer withholds the whole annotation. One bad gene in a 60,000-gene genome used to produce exit 2 and only a
*.partial.gff3. Per-locus failures now publish, reportpartial_success, and are recorded inrun_manifest.json;--strict-completenessrestores the old behaviour.Write-funnel crashes on gene-like/organellar children and on inverted (
start > end) coordinates, which aborted a whole-genome write.Three-level hierarchies (
gene → primary_transcript → miRNA → exon) are lifted instead of silently dropped, and LiftOn no longer rejects its own valid Ensembl/GENCODE-shaped output.
Added:
Auditable run manifests and transactional output —
lifton_output/run_manifest.jsonrecords sanitized arguments, SHA-256 input fingerprints, tool versions, phase timings, counts, validation and failures; the GFF3 is staged and published atomically only on success.Always-on structural output validation before publication, recorded in the run manifest and not bypassable by
--allow-partial-output.Content-addressed annotation caches that rebuild when their manifest no longer matches the source, parser settings, backend, or LiftOn version.
Performance (byte-neutral):
GFF3 validation is bounded — one contiguous top-level block at a time instead of the whole file: 5.64 GB → 267 MB on a full dog→cat lift (21× less memory, 1.9× faster), with identical reports.
Step 7 is ~21% faster on Drosophila at
-t 8(redundant leaf queries removed, vectorized identity counters, a pruning ordered gate), with byte-identical output and flat peak memory.Features are cloned rather than deep-copied (33.06 µs → 5.00 µs per rich CDS), and per-stage in-flight bounds (
--step7-max-inflightand siblings) cap the memory a parallel run can hold.
v1.0.9
Incremental release (2026-06-21). Turns on several accuracy- and
completeness-improving defaults, hardens LiftOn against whole-genome-abort
crashes on full RefSeq / cross-species genomes, and adds byte-identical
performance fast-paths plus validation tools. Some new defaults change the
output annotation relative to v1.0.8 — each ships with an opt-out flag that
restores the previous behaviour. See the project CHANGELOG.md for the full
list.
Changed (output-affecting defaults — results differ vs v1.0.8):
Gene-like lift is now the default. LiftOn auto-detects every reference top-level parent type with a transcript/exon hierarchy and lifts them all (pseudogenes,
ncRNA_gene, structured mobile elements, ...), not justgene. This adds features. Pass--gene-onlyto restore the oldgene-only lift (--lift-gene-likeis a kept no-op alias).Best-of-outcome Liftoff/miniprot merge is now the default. Per transcript, LiftOn keeps whichever of {chained merge + ORF-rescue, Liftoff + ORF-rescue} yields the higher emitted protein identity, avoiding merges that could silently frameshift downstream CDS. Pass
--legacy-mergeto restore the pre-promotion unconditional merge (--optimizeis a kept no-op alias).Banded / windowed alignment is now the default for all gene sizes. The aligner uses anchor-windowed alignment above ~2500 aa / 8000 nt (giant genes are always memory-bounded, so titin-scale transcripts no longer OOM): identity-exact on same-species lifts, mean-neutral cross-species, much faster and lighter on memory. Pass
--full-dp-alignto restore the exact giant-only full-DP path (--fast-alignis a kept no-op alias).Miniprot-only rescue is now default-ON. When the DNA lift misses a reference coding gene entirely (its miniprot mRNA overlaps no lifted gene locus), LiftOn emits the miniprot-only model, tagged
lifton_rescue=miniprot_only. It runs as a separate pass after the main lift closes, gated by a protein-identity floor with a dedup guard, recovering genuinely-missing genes at large evolutionary distance. Pass--no-miniprot-rescue(orLIFTON_MINIPROT_RESCUE=0) to opt out (--miniprot-rescueis a kept no-op alias).
Fixed (robustness / crashes):
Gene-like child double-lift crash on full RefSeq genomes. A gene-like feature that is a child of a gene was being enumerated again as a top-level locus, producing a duplicate FASTA key and crashing the run. Full Arabidopsis now completes (~99.9% of coding transcripts recovered, vs ~28% before the crash) and full rice likewise (~77% → ~99.9%). Top-level-only annotations are byte-identical.
Inverted-coordinate write crash. A single malformed transcript with
start > end(e.g. on the dog→cat lift) used to abort an entire ~60k-transcript genome during the write phase; such a feature is now skipped and logged, and the rest of the genome completes.A malformed feature no longer aborts a whole genome. The transcript writer now catches the project's validation exception so one bad feature is dropped and logged instead of propagating out of the parent write phase.
Whole-genome-abort hardening around
consume()/__str__plus a recursion-limit guard and a full traceback dump in the vendored-Liftoff call path, so deep/odd inputs fail loudly per-feature instead of silently killing the run.
Added (performance — byte-identical fast-paths, same output, faster/lighter):
--stream— pipe miniprot output straight into an in-memory database, skipping theminiprot.gff3disk round-trip and SQLite re-ingest.--inmemory-liftoff— feed Liftoff's lifted features to the database in-process, skipping theliftoff.gff3disk write and re-ingest.--threads N --locus-pipeline— fan out per-locus work across a thread pool; output is emitted in submission order so--threads Nis byte-identical to--threads 1(now works on the default backend without--native).--native— enable experimental native compatibility hooks. The mappy Liftoff route also requiresLIFTON_NATIVE_LIFTOFF_ALIGN=1; miniprot and bounded locus workers keep their proven paths.Concurrent aligner step is now the default (miniprot and Liftoff overlap); pass
--serial-alignersto opt out (--parallel-alignersis a kept no-op alias).Fused parallel Step 7 — per-locus materialise and process phases are fused into one pool, lowering both wall-clock and peak memory.
Sequence-extraction query collapse — collapses per-feature database queries (~2.2–2.4× fewer round-trips, ~25–34% faster extraction).
miniprot ``-t`` now scales with ``--threads`` (was pinned to its built-in default of 4; the default
-t 1is byte-identical).
Added (validation):
--strict-gff— run the NCBI GFF3 input-side validator on the reference annotation and exit non-zero on any spec violation.--validate-output/--validate-verbose— re-validate the just-written output GFF3 and print a structured report.``gff3-validate`` console script — a standalone GFF3 validator installed alongside
lifton.
Packaging:
Python floor raised to
>=3.10(thenetworkx>=3.3dependency requires Python ≥3.10; 3.9 is EOL).mappyis an optional dependency for the explicitly enabled native Liftoff route; the runtime falls back gracefully when it is absent.Added
MANIFEST.in,pyproject.toml(PEP 517/518 build config), and PyPI trove classifiers / project URLs.The vendored
gffbaseships its pure-Python fallback parser (no pre-built.so), so installs work without a Rust toolchain.
v1.0.7
Bug Fixes:
Fixed gffutils UNIQUE constraint errors: Enhanced duplicate feature ID handling with automatic recovery strategies. The system now automatically handles duplicate IDs in polished liftoff output and RefSeq annotations with non-overlapping CDS features, preventing silent failures during database creation.
Fixed GTF file format processing: Added automatic GTF format detection and proper handling. GTF files are now correctly processed with automatic gene/transcript inference, and optional automatic conversion to GFF3 format using gffread or agat tools.
Fixed ID parsing for IDs ending with numbers: Improved get_ID_base() function to safely handle feature IDs that naturally end with underscore and number (e.g., FMUND_1). The function now only removes suffixes when confirmed to be copy numbers, preventing silent failures.
Fixed CDS ID preservation: CDS features now preserve their IDs in output files, complying with GFF3 specification. CDS features from the same mRNA can now share the same ID as required by the standard.
Improvements:
Enhanced biotype attribute support: Added support for generic biotype attribute as fallback when gene_biotype (RefSeq) or gene_type (GENCODE/ENSEMBL/CHESS) are not present. This ensures protein-coding features are correctly identified regardless of annotation source.
Automatic GTF to GFF3 conversion: Added optional automatic conversion of GTF files to GFF3 format for better compatibility. Conversion uses gffread (preferred) or agat tools if available, with graceful fallback to direct GTF processing.
Improved error handling: Enhanced error messages and recovery strategies for database creation failures, providing better user guidance and automatic problem resolution.
Better format detection: Improved file format detection logic that checks multiple lines and patterns to reliably distinguish between GTF and GFF3 formats, with GFF3 as the safe default.
v1.0.0
Initial release of LiftOn
Release via the documentation (https://khchao.com/LiftOn)
Released via the paper (bioRxiv coming soon!)