pathotypr predict¶
Assign a lineage to each genome in a FASTA file using a model produced by pathotypr train.
pathotypr predict classifies query sequences against a pre-trained random-forest model bundle. Each record receives a majority-vote lineage call together with a confidence score, a confidence margin, and the runner-up votes. Use it once you have a trained .pathotypr.zst model and want to genotype new assemblies or genomes — the input FASTA is streamed in batches, so memory stays constant even for very large collections.
Synopsis¶
Inputs¶
predict needs two files.
| What | Flag | Format |
|---|---|---|
| Genomes to classify | -i, --input |
FASTA, plain or gzipped. Multi-record is fine, and each record is classified separately |
| Trained model | -m, --model |
A .pathotypr.zst bundle written by train |
Requirements¶
- [x] The model must be a current-format bundle. An incompatible format version is a hard error and the model has to be retrained. A bundle written by a different pathotypr version only warns.
- [x] Records shorter than the model's k are skipped with a warning: no k-mer can be extracted from them.
- [x] No labels are needed or read. The full header is carried through to the output.
The model dictates k, not you
predict has no -k. The k-mer size is whatever train used and is stored
in the bundle, so a model trained at k=21 always vectorises queries at
k=21. This is also why the minimum usable sequence length depends on the
model rather than on the query file.
Every record gets a call
There is no "unknown" class and no confidence threshold. A query from a lineage absent from the training set still receives the closest label the forest can find, distinguishable only by a low confidence and margin. Read those two columns before acting on a call: see Output for what they mean.
Options¶
| Option | Default | Required | Description |
|---|---|---|---|
-i, --input <FILE> |
— | yes | Input FASTA of sequences to classify (gzip supported). |
-m, --model <FILE> |
— | yes | Model bundle produced by pathotypr train. |
-o, --output <FILE> |
— | yes | Output predictions TSV. |
-t, --threads <N> |
all cores | no | Number of CPU threads. |
--excel |
false |
no | Also write an Excel (.xlsx) file alongside the TSV. |
Global flags¶
These are available on every pathotypr subcommand:
| Flag | Description |
|---|---|
-v, --verbose |
Increase log verbosity. Repeatable: -v = debug, -vv = trace. |
-h, --help |
Print help. |
The model carries its own settings
The k-mer size and feature-hashing parameters are stored inside the model bundle. You do not (and cannot) set them at predict time — the same vectorization used during training is applied automatically, so predictions stay consistent with the model.
How it works¶
- Load the model bundle and decompress it (zstd).
- Validate the model bundle: the run stops only if its format version differs from the one this build expects. A model trained by a different
pathotyprrelease with the same format loads fine, with a warning. - Stream the input FASTA in batches of 512 records — RAM usage is proportional to the batch, not the number of sequences.
- Vectorize each batch with the exact feature hasher stored in the model.
- Predict in parallel: every tree in the forest votes, and the majority class becomes the call.
- Score each record with a confidence and a confidence margin, then write results immediately and free the batch.
Short sequences are skipped
Any record shorter than the model's k-mer length yields no k-mers and is skipped with a warning; it will not appear in the output. If every record is skipped or the file is empty, the command exits with an error and removes the partial output.
Examples¶
Reading confidence
Confidence is the fraction of trees that voted for the winning lineage (e.g. 0.9500 means 95% of trees agreed). Confidence_Margin is how far ahead the winner is over the runner-up (e.g. 0.4000 means the winner led the second-best class by 40% of the trees). A high confidence with a low margin points to two closely competing lineages — inspect Other_Votes.
Output¶
predictions.tsv (always written)¶
A tab-separated file with a header row and one row per classified record. Columns are exactly:
| Column | Description |
|---|---|
Header |
The input record's full FASTA header line — everything after >, description included. |
Predicted_Lineage |
Majority-vote lineage label. A record on which no tree cast a usable vote still gets a label, but with a Confidence at or near zero — read low confidence, not a special value, as "unresolved". |
Confidence |
Fraction of trees voting for the winning lineage, in 0.0000–1.0000 (4 decimals). |
Confidence_Margin |
Winner-minus-runner-up vote fraction, in 0.0000–1.0000 (4 decimals). |
Other_Votes |
Up to the three next-best lineages as label:fraction (2 decimals), comma-separated. Empty when there are no alternatives. |
Example:
Header Predicted_Lineage Confidence Confidence_Margin Other_Votes
sample_001 L4 0.9600 0.9300 L2:0.03,L1:0.01
sample_002 L2 0.5400 0.0800 L4:0.46
predictions.xlsx (only with --excel)¶
When --excel is set, an Excel workbook with the same five columns is written next to the TSV (the output path with an .xlsx extension). Conditional formatting is applied to the Confidence and Confidence_Margin columns to highlight lower-scoring calls at a glance.
Excel is best-effort
The TSV is the authoritative output. If the Excel writer cannot be initialized or a row fails to append, predict logs a warning and continues writing the TSV normally.
See also¶
- Train — build the
.pathotypr.zstmodel bundle used here. - Input formats — accepted FASTA/gzip inputs.
- Classify — marker-based assembly classification.
- Prediction algorithm — streaming batches, ensemble voting, confidence and margin metrics.
- Random forest — how each tree votes.
- Feature hashing — how sequences are vectorized identically at train and predict time.