Choose your input and output formats. The form will explain any special requirements before conversion.
Available formats
- ABI / AB1 trace — Sanger capillary sequencing trace (ABI/AB1). Read-only input: the base calls and their quality scores are extracted.
- ABI / AB1 trace (quality trimmed) — ABI/AB1 trace with Mott's quality trimming applied to the read ends. Read-only input.
- ACE assembly contigs — Sequence assembly file. Read-only input: contig sequences and their quality scores are extracted.
- mmCIF structure (observed residues) — mmCIF structure. Read-only input: the protein sequence is built from observed residues in the atom coordinates.
- mmCIF structure (declared sequence) — mmCIF structure. Read-only input: the full declared polymer sequence is read.
- Clustal alignment — Clustal multiple sequence alignments.
- EMBL — EMBL flat-file sequence records with features and annotations.
- EMBL coding sequences (CDS) — EMBL flat file. Read-only input: extracts CDS features as protein records, one per coding sequence.
- EMBOSS pairwise alignment — EMBOSS pairwise alignment output, for example from needle or water. Read-only input.
- FASTA — FASTA sequence records.
- FASTA (two-line) — FASTA with exactly two lines per sequence record.
- FASTA (BLAST-style comments) — FASTA that may contain comment lines starting with !, #, or ;. Read-only input.
- FASTA -m 10 alignment output — Pairwise alignments from the FASTA package's -m 10 output. Read-only input.
- FASTA (Pearson-style comments) — FASTA that may have comment lines before the first record or lines starting with ;, as in William Pearson's FASTA software. Read-only input.
- FASTQ (Sanger PHRED+33) — Sanger FASTQ records with PHRED quality scores.
- FASTQ (Illumina PHRED+64) — FASTQ records using Illumina quality scores.
- FASTQ (Solexa+64) — FASTQ records using Solexa quality scores.
- Gene Construction Kit (GCK) — Gene Construction Kit file holding one sequence. Read-only input.
- GenBank — GenBank or GenPept flat-file sequence records.
- GenBank coding sequences (CDS) — GenBank flat file. Read-only input: extracts CDS features as protein records, one per coding sequence.
- GFA 1 graph segments — Graphical Fragment Assembly version 1 graph. Read-only input: segment sequences are extracted.
- GFA 2 graph segments — Graphical Fragment Assembly version 2 graph. Read-only input: segment sequences are extracted.
- IntelliGenetics — IntelliGenetics sequence files. Read-only input.
- IMGT (EMBL variant) — IMGT variant of the EMBL flat-file format.
- MAF alignment — Multiple Alignment Format (MAF) sequence alignments.
- Mauve XMFA alignment — Mauve extended multi-FASTA (XMFA) alignments.
- MSF alignment — MSF multiple sequence alignments. Read-only input.
- NEXUS alignment — NEXUS multiple sequence alignments.
- UCSC nib — UCSC nib file holding one DNA sequence (A, C, G, T, N) without a record name.
- PDB structure (observed residues) — PDB structure. Read-only input: the protein sequence is built from observed residues in the atom coordinates.
- PDB structure (declared sequence) — PDB structure. Read-only input: the protein sequence declared in the SEQRES records is read.
- PHRED PHD — PHRED PHD base calls with per-base quality scores.
- PHYLIP alignment (interleaved) — Interleaved PHYLIP alignments with IDs of up to 10 characters.
- PHYLIP alignment (relaxed names) — PHYLIP alignments with longer IDs that contain no spaces.
- PHYLIP alignment (sequential) — Sequential PHYLIP alignments with IDs of up to 10 characters.
- PIR / NBRF — PIR / NBRF sequence records.
- QUAL quality scores — Quality scores in a FASTA-like layout, without sequence letters.
- SeqXML — Simple Sequence XML records.
- SFF flowgram reads — Standard Flowgram Format reads with quality scores and flow data.
- SFF flowgram reads (trimmed) — SFF reads with the trimming recorded in the file applied. Read-only input.
- SnapGene — SnapGene file holding one sequence with its features. Read-only input.
- Stockholm alignment — Stockholm (Pfam) multiple sequence alignments.
- Swiss-Prot / UniProt text — Swiss-Prot / UniProt flat-file protein records. Read-only input.
- Tab-separated (ID and sequence) — Two-column tab-separated records: ID, then sequence.
- UCSC twoBit — UCSC twoBit genome sequence files. Read-only input.
- UniProt XML — UniProt XML protein records. Read-only input.
- DNA Strider / Serial Cloner (xdna) — DNA Strider / Serial Cloner file holding one sequence.
Sample input file
Choose an input format above to see a small example file you can copy, paste, or download to try the converter.
Samples are small illustrative files. Some come from the Biopython test suite (see the project's sample notes); they are not real study data.