nf-core/proteinfamilies
Generation and updating of protein families
Define where the pipeline should find input data and save output data.
Path to comma-separated file ‘.csv’ containing information about the samples in the experiment.
string^\S+\.csv$The output directory where the results will be saved. You have to use absolute paths to storage on Cloud infrastructure.
stringEmail address for completion summary. Example: name.surname@example.com
string^([a-zA-Z0-9_\-\.]+)@([a-zA-Z0-9_\-\.]+)\.([a-zA-Z]{2,5})$MultiQC report title. Printed as page header, used for filename if not otherwise specified.
stringPublish intermediate files under intermediates/ (e.g. clustering databases, raw alignments and HMMs, search results).
booleanParameters used to describe centralised config profiles. These should not be edited.
Git commit id for Institutional configs.
stringmasterBase directory for Institutional configs.
stringhttps://raw.githubusercontent.com/nf-core/configs/masterInstitutional config name.
stringInstitutional config description.
stringInstitutional config contact information.
stringInstitutional config URL link.
stringBase path / URL for data used in the modules
stringLess common options for the pipeline, typically set in a config file.
Display version and exit.
booleanMethod used to save pipeline results to output directory.
stringEmail address for completion summary, only when pipeline fails. Example: name.surname@example.com
string^([a-zA-Z0-9_\-\.]+)@([a-zA-Z0-9_\-\.]+)\.([a-zA-Z]{2,5})$Send plain-text email instead of HTML.
booleanFile size limit when attaching MultiQC reports to summary emails. Example: name.surname@example.com
string25.MB^\d+(\.\d+)?\.?\s*(K|M|G|T)?B$Do not use coloured log outputs.
booleanCustom config file to supply to MultiQC.
stringCustom logo file to supply to MultiQC. File name must also be set in the MultiQC config file
stringCustom MultiQC yaml file containing HTML including a methods description.
stringBoolean whether to validate parameters against the schema at runtime
booleantrueBase URL or local path to location of pipeline test dataset files
stringhttps://raw.githubusercontent.com/nf-core/test-datasets/proteinfamilies/Suffix to add to the trace report filename. Default is the date and time in the format yyyy-MM-dd_HH-mm-ss.
stringDisplay the help message.
boolean,stringDisplay the full detailed help message.
booleanDisplay hidden parameters in the help message (only works when –help or –help_full are provided).
booleanUse these parameters to control the quality check and preprocessing of the input sequences.
Skip all default QC steps for sequences (gap trimming, length filtering, validation, duplicate removal).
booleanThe minimum allowed sequence length
integer30The maximum allowed sequence length
integer5000How duplicate input sequences are found and removed.
stringUse these parameters to control the initial clustering of the input sequences.
Choose clustering algorithm. Either simple ‘cluster’ for medium size inputs, or ‘linclust’ for less sensitive clustering of larger datasets.
stringmmseqs parameter for minimum sequence identity
number0.3mmseqs parameter for minimum sequence coverage ratio
number0.5mmseqs parameter for coverage mode: 0 for both, 1 for target and 2 for query sequence
integerMinimum clustering chunk size threshold to create seed Multiple Sequence Alignments upon.
integer25Use these parameters to choose and control the family generation algorithm, and the update of existing families.
Choose the algorithm that turns clusters into family models. Either ‘standard’, aligning each cluster and building one HMM per task, or ‘iterative’, letting mgnifam loop HMM building, recruitment and realignment per cluster.
stringNumber of clusters handed to each family generation task by the ‘iterative’ algorithm.
integer1000Choose alignment tool. FAMSA is recommended as best time-memory-accuracy combination option.
stringKeep the existing HMMs of updated families, only rebuilding their full MSAs.
booleanUse these parameters to control the trimming of the seed multiple sequence alignments.
Skip trimming gappy positions from seed Multiple Sequence Alignments (MSAs). Full MSAs are never trimmed.
booleanChoose if ClipKIT should only clip gaps at the ends of the MSAs.
booleantrueMultiple Sequence Alignment (MSA) positions with gappiness greater than this threshold will be trimmed
number0.5Use these parameters to control the HMM searches and the recruitment of additional sequences into the families.
Skip recruitment of additional sequences from the input FASTA file using the family Hidden Markov Models (HMMs) into the full alignment
booleanhmmsearch e-value cutoff threshold for reported results
number0.001hmmsearch minimum length percentage filter of hit env vs query length
number0.9Use these parameters to control the removal of redundant families and the merging of similar families.
Which families take part in the removal of between-family redundancy (hmmsearch of family representatives against family models).
stringWhich families can be merged with similar (but not redundant) families.
stringName of a merged family that holds updated families.
stringhmmsearch minimum length percentage filter of hit env vs query length, for redundant family removal
number1hmmsearch minimum length percentage of hit env vs query length, to flag and report similar families (and to optionally merge)
number0.9Use these parameters to control the removal of redundant sequences within each family.
Skip removal of inside-family redundancy of sequences via mmseqs clustering.
booleanmmseqs parameter for minimum sequence identity
number0.9mmseqs parameter for minimum sequence coverage ratio
number0.9mmseqs parameter for coverage mode: 0 for both, 1 for target and 2 for query sequence
integerUse these parameters to enable the phylogenetic tree construction of the final full MSAs.
Infer a phylogenetic tree of each final family’s full MSA with CMAPLE.
boolean