User Manual

LiftOn

====================================================================
An accurate homology lift-over tool between assemblies
====================================================================


   ██╗     ██╗███████╗████████╗ ██████╗ ███╗   ██╗
   ██║     ██║██╔════╝╚══██╔══╝██╔═══██╗████╗  ██║
   ██║     ██║█████╗     ██║   ██║   ██║██╔██╗ ██║
   ██║     ██║██╔══╝     ██║   ██║   ██║██║╚██╗██║
   ███████╗██║██║        ██║   ╚██████╔╝██║ ╚████║
   ╚══════╝╚═╝╚═╝        ╚═╝    ╚═════╝ ╚═╝  ╚═══╝

v1.0.10

usage: lifton [-h] [-E] [-EL] [-c] [--no-orf-search] [-o FILE] [-u FILE] [-exclude_partial]
            [-mm2_options =STR] [-mp_options =STR] [-a A] [-s S] [-min_miniprot MIN_MINIPROT]
            [-max_miniprot MAX_MINIPROT] [-d D] [-flank F] [-V] [-D] [-t THREADS] [-m PATH]
            [-f TYPES] [-infer-genes] [-infer_transcripts] [-chroms TXT] [-unplaced TXT]
            [-copies] [-sc SC] [-overlap O] [-mismatch M] [-gap_open GO] [-gap_extend GE]
            [-polish] [-cds] [-time] [--validate-output] [--validate-verbose]
            [--allow-partial-output] [--strict-completeness] [--strict-gff] [--stream]
            [--inmemory-liftoff] [--locus-pipeline] [--step7-max-inflight N]
            [--step8-max-inflight N] [--evaluation-max-inflight N] [--native]
            [--serial-aligners] [--parallel-aligners] [--optimize] [--legacy-merge]
            [--full-dp-align] [--fast-align] [--gene-only] [--lift-gene-like]
            [--no-miniprot-rescue] [--miniprot-rescue] [--miniprot-cross-locus-rescue]
            [--no-miniprot-candidate] [--miniprot-candidate] [--no-adaptive-rescue-floor]
            [--adaptive-rescue-floor] [--merge-strategy STRATEGY] [--id-spec ID_SPEC] [--force]
            [--verbose] [--no-auto-convert-gtf] -g GFF [-P FASTA] [-T FASTA] [-L gff] [-M gff]
            [-ad SOURCE]
            target reference

Lift features from one genome assembly to another

* Required input (sequences):
target                target fasta genome to lift genes to
reference             reference fasta genome to lift genes from

* Required input (Reference annotation):
-g GFF, --reference-annotation GFF
                        the reference annotation file to lift over in GFF or GTF format (or) name of feature database; if not specified, the -g argument must be
                        provided and a database will be built automatically

* Optional input (Reference sequences):
-P FASTA, --proteins FASTA
                        the reference protein sequences.
-T FASTA, --transcripts FASTA
                        the reference transcript sequences.

* Optional input (Liftoff annotation):
-L gff, --liftoff gff
                        the annotation generated by Liftoff (or) name of Liftoff gffutils database; if not specified, the -liftoff argument must be provided and a
                        database will be built automatically

* Optional input (miniprot annotation):
-M gff, --miniprot gff
                        the annotation generated by miniprot (or) name of miniprot gffutils database; if not specified, the -miniprot argument must be provided
                        and a database will be built automatically

* gffutils parameters:
--merge-strategy {create_unique,merge,error,warning,replace}
                        strategy for merging features when building the database; default "create_unique"
--id-spec ID_SPEC     attribute to use as feature ID; default "ID"
--force               overwrite existing database
--verbose             enable verbose output

* Output settings:
-o FILE, --output FILE
                        write output to FILE in same format as input; by default, output is written to "lifton.gff3"
-u FILE               write unmapped features to FILE; default is "unmapped_features.txt"
-exclude_partial      write partial mappings below -s and -a threshold to unmapped_features.txt; if true partial/low sequence identity mappings will be included
                        in the gff file with partial_mapping=True, low_identity=True in comments
--allow-partial-output
                        publish structurally valid output despite recorded
                        per-locus failures; otherwise preserve it as
                        *.partial.gff3 and exit non-zero

* Miscellaneous settings:
-h, --help            show this help message and exit
-E, --evaluation      Run LiftOn in evaluation mode
-EL, --evaluation-liftoff-chm13
                        Run LiftOn in evaluation mode
-c, --write_chains    Write chaining files
--no-orf-search       do not perform open reading frame (ORF) search
-V, --version         show program version
-D, --debug           Run debug mode
-t THREADS, --threads THREADS
                        use t parallel processes to accelerate alignment; by default t=1
-m PATH               Minimap2 path
-f TYPES, --features TYPES
                        list of feature types to lift over (an explicit -f overrides gene-like auto-detection)
-infer-genes          use if annotation file only includes transcripts, exon/CDS features; auto-enabled for GTF
-infer_transcripts    use if annotation file only includes exon/CDS features and does not include transcripts/mRNA; auto-enabled for GTF
-chroms TXT           comma seperated file with corresponding chromosomes in the reference,target sequences
-unplaced TXT         text file with name(s) of unplaced sequences to map genes from after genes from chromosomes in chroms.txt are mapped; default is
                        "unplaced_seq_names.txt"
-copies               look for extra gene copies in the target genome
-sc SC                with -copies, minimum sequence identity in exons/CDS for which a gene is considered a copy; must be greater than -s; default is 1.0
-overlap O            maximum fraction [0.0-1.0] of overlap allowed by 2 features; by default O=0.1
-mismatch M           mismatch penalty in exons when finding best mapping; by default M=2
-gap_open GO          gap open penalty in exons when finding best mapping; by default GO=2
-gap_extend GE        gap extend penalty in exons when finding best mapping; by default GE=1
-polish
-cds                  annotate status of each CDS (partial, missing start, missing stop, inframe stop codon)
-time, --measure_time enable time measurement for each step (writes time.txt)
-ad SOURCE, --annotation-database SOURCE
                        The source of the reference annotation (RefSeq / GENCODE / others)
--no-auto-convert-gtf disable automatic GTF -> GFF3 conversion

* Validation (does not change output bytes):
(always on)           stream-validate GFF3 structure before atomic
                      publication; failures preserve *.partial.gff3 and
                      exit non-zero
--strict-gff          run the NCBI GFF3 input-side validator on the reference annotation; exit non-zero on any spec violation
--validate-output     add full hierarchy, containment, phase, and LiftOn-attribute validation before publication
--validate-verbose    with --validate-output, also print warnings (not just errors)

* Performance fast-paths (BYTE-IDENTICAL to the default output):
--stream              incrementally ingest miniprot stdout into a staged
                      DuckDB database; skip miniprot.gff3 and duplicate
                      graph materialization
--inmemory-liftoff    feed Liftoff's lifted features to the database in-process; skip the liftoff.gff3 disk write / re-ingest
--locus-pipeline      fan out Steps 7, 8, and evaluation through bounded
                      workers sized by --threads; publish in submission
                      order (byte-identical to -t 1)
--step7-max-inflight N
                        bound submitted-but-not-emitted Step-7 loci; default 2 * --threads
--step8-max-inflight N
                        bound analyzed-but-not-published Step-8 candidates; default 2 * --threads
--evaluation-max-inflight N
                        bound submitted-but-not-written evaluation loci; default 2 * --threads
--native              enable experimental native compatibility hooks; combine with LIFTON_NATIVE_LIFTOFF_ALIGN=1 to opt into mappy, while miniprot and bounded locus workers retain their proven paths
--serial-aligners     opt OUT of the (default) concurrent Liftoff||miniprot overlap; run them sequentially

* Output-changing flags (opt-outs that RESTORE earlier behaviour):
--legacy-merge        restore the pre-promotion UNCONDITIONAL Liftoff/miniprot merge (default = verified best-of-outcome merge)
--full-dp-align       restore the exact giant-only full-DP alignment (default = band-everything anchor-windowed alignment)
--gene-only           restore the pre-v1.0.9 gene-only lift (default = lift all auto-detected gene-like top-level types)
--no-miniprot-rescue  disable the (default-ON) miniprot-only rescue pass for genes the DNA lift missed entirely
--no-miniprot-candidate
                        opt OUT of the (default-ON) 3rd best-of-outcome candidate: miniprot's NATIVE CDS-only model, kept only when its ORF-rescued protein identity is STRICTLY better than
                        the 2-way winner, so per-transcript identity is non-decreasing. Env LIFTON_MINIPROT_CANDIDATE=0 also disables it
--no-adaptive-rescue-floor
                        restore the FIXED 0.50 miniprot-only-rescue floor (default = lower the floor toward 0.30 as the DNA lift's gene recall drops, inert on same/close-species). Env
                        LIFTON_RESCUE_ADAPTIVE_FLOOR=0 also disables it

* Experimental, opt-in:
--miniprot-cross-locus-rescue
                        replace a WEAKLY lifted coding gene (best emitted protein identity < LIFTON_CROSS_LOCUS_MAX_LIFTOFF, default 0.5) with a clean miniprot model on a DIFFERENT chromosome,
                        tagged lifton_rescue=cross_locus. Off by default

* Completeness / publication:
--strict-completeness
                        refuse to publish if ANY locus was skipped (default = publish, report the run as "partial_success", and record every skipped locus in run_manifest.json)
--allow-partial-output
                        publish the staged output even when a failure would otherwise block publication

* Deprecated NO-OP aliases (kept for backward compatibility; have no effect):
--parallel-aligners   no-op (concurrent Step 4 is now default; use --serial-aligners to opt out)
--optimize            no-op (best-of-outcome merge is now default; use --legacy-merge to opt out)
--fast-align          no-op (band-everything alignment is now default; use --full-dp-align to opt out)
--lift-gene-like      no-op (gene-like lift is now default; use --gene-only to opt out)
--miniprot-rescue     no-op (miniprot-only rescue is now default; use --no-miniprot-rescue to opt out)
--miniprot-candidate  no-op (the miniprot-only candidate is now default; use --no-miniprot-candidate to opt out)
--adaptive-rescue-floor
                        no-op (the adaptive rescue floor is now default; use --no-adaptive-rescue-floor to opt out)

Alignments:
-mm2_options =STR     space delimited minimap2 parameters. By default ="-a --end-bonus 5 --eqx -N 50 -p 0.5"
-mp_options =STR      space delimited miniprot parameters. By default ""
-a A                  designate a feature mapped only if it aligns with coverage ≥A; by default A=0.5
-s S                  designate a feature mapped only if its child features (usually exons/CDS) align with sequence identity ≥S; by default S=0.5
-min_miniprot MIN_MINIPROT
                        The minimum length ratio of a protein-coding transcript to the longest protein-coding transcript within a gene locus, as identified
                        exclusively by miniprot in the target genome, is set by default to MIN_MINIPROT=0.9.
-max_miniprot MAX_MINIPROT
                        The maximum length ratio of a protein-coding transcript to the longest protein-coding transcript within a gene locus, as identified
                        exclusively by miniprot in the target genome, is set by default to MAX_MINIPROT=1.5.
-d D                  distance scaling factor; alignment nodes separated by more than a factor of D in the target genome will not be connected in the graph; by
                        default D=2.0
-flank F              amount of flanking sequence to align as a fraction [0.0-1.0] of gene length. This can improve gene alignment where gene structure differs
                        between target and reference; by default F=0.0

Output-changing defaults and their opt-outs

Several defaults toggle behaviour that was introduced or promoted after v1.0.8. The tables below summarise, for each flag, whether it CHANGES the output annotation or is BYTE-NEUTRAL (same output bytes, different I/O or scheduling), and which flag restores the older behaviour.

Output-changing defaults (opt-out restores the older behaviour):

Default behaviour

Opt-out flag

Effect

Gene-like lift (lifts pseudogenes / ncRNA_gene / mobile elements, not just gene) — v1.0.9

--gene-only

CHANGES output (adds features). --lift-gene-like is a deprecated no-op alias.

Best-of-outcome Liftoff/miniprot merge — v1.0.9

--legacy-merge

CHANGES output. Keeps the higher-identity of {merge, Liftoff} per transcript. --optimize is a deprecated no-op alias.

Band-everything / windowed alignment — v1.0.9

--full-dp-align

CHANGES output on divergent inputs only (identity-exact same-species). --fast-align is a deprecated no-op alias.

Miniprot-only rescue — v1.0.9

--no-miniprot-rescue

CHANGES output (adds lifton_rescue=miniprot_only genes the DNA lift missed). --miniprot-rescue is a deprecated no-op alias. Env LIFTON_MINIPROT_RESCUE=0/1 force-disables/enables.

Third merge candidate: miniprot's native CDS-only model — v1.0.10

--no-miniprot-candidate

CHANGES output, non-decreasing per transcript: adopted only when its ORF-rescued protein identity is STRICTLY better than the two-way winner, and never for an antisense hit. --miniprot-candidate is a deprecated no-op alias. Env LIFTON_MINIPROT_CANDIDATE=0/1.

Divergence-adaptive miniprot-only rescue floor — v1.0.10

--no-adaptive-rescue-floor

CHANGES output (adds genes at large evolutionary distance). Lowers the 0.50 identity floor toward 0.30 as the DNA lift's gene recall drops; inert on same/close-species. --adaptive-rescue-floor is a deprecated no-op alias. Env LIFTON_RESCUE_ADAPTIVE_FLOOR=0/1.

Coding transcripts harmonized to mRNA; every CDS carries an ID and the reference's descriptive attributes — v1.0.10

LIFTON_NO_MRNA_HARMONIZE=1 / LIFTON_NO_CDS_ATTR_CARRY=1 / LIFTON_NO_CONTAINMENT_NORMALIZE=1

CHANGES column 3 and column 9 only; coordinates and the encoded protein are untouched. Output grows 12–43%.

Byte-neutral performance flags (identical output, faster / lighter):

Flag

Effect (BYTE-IDENTICAL to the default output)

--stream

Incrementally ingest miniprot stdout into a staged, checkpointed DuckDB database; skip miniprot.gff3 and duplicate graph materialization.

--inmemory-liftoff

Feed Liftoff's lifted features to the database in-process; skip the liftoff.gff3 disk write and re-ingest.

--threads N --locus-pipeline

Fan out Steps 7, 8, and evaluation; ordered publication keeps --threads N == --threads 1 byte-for-byte. Each stage defaults to at most 2 * N in-flight items; tune it with --step7-max-inflight, --step8-max-inflight, or --evaluation-max-inflight to trade utilization for memory.

--native

Enable experimental native compatibility hooks. With LIFTON_NATIVE_LIFTOFF_ALIGN=1, opt into mappy for Liftoff; miniprot keeps its guarded subprocess/direct-stream path and bounded locus workers do not require this flag.

--serial-aligners

Opt out of the (default) concurrent Liftoff/miniprot overlap; useful on core-constrained machines. --parallel-aligners is a deprecated no-op alias.

Validation (changes the exit code, never output bytes): Every staged output must pass structural validation before atomic publication. This gate rejects malformed records and hierarchy references even without a flag, and --allow-partial-output cannot bypass it. --strict-gff validates the reference input; --validate-output adds the full output validator, with --validate-verbose showing warnings. The standalone gff3-validate console script is installed alongside lifton.