Scanpy is a scalable toolkit for analysing single-cell gene expression data.
It provides facilities for preprocessing, visualisation, clustering, trajectory inference and differential expression testing.
Built jointly with AnnData, Scanpy can efficiently handle datasets containing more than one million cells. Experimental Dask support lets some functions process datasets too large to fit in memory.
This is free and open source software.
Key Features
- Calculates quality-control metrics and filters cells and genes.
- Normalises, transforms, regresses and scales expression data.
- Identifies highly variable genes.
- Offers dimensionality reduction with PCA, t-SNE, UMAP and diffusion maps.
- Computes neighbourhood graphs for single-cell datasets.
- Supports Leiden and hierarchical clustering.
- Provides trajectory inference using PAGA and diffusion pseudotime.
- Ranks marker genes and performs differential expression testing.
- Creates scatter plots, heatmaps, dot plots, violin plots and matrix plots.
- Integrates with BBKNN, Harmony, MNN and Scanorama.
- Uses AnnData for storing annotated data matrices.
- Offers experimental Dask support for larger-than-memory datasets.
Website: github.com/scverse/scanpy
Support:
Developer: scverse
License: BSD 3-Clause “New” or “Revised” License
Scanpy is written in Python. Learn Python with our recommended free books and free tutorials.
Related Software
| Bioinformatics Tools | |
|---|---|
| Bioconductor | Analysis and comprehension of high-throughput genomic data |
| Biopython | Tools for biological computation written in Python |
| UGENE | Set of integrated bioinformatics software |
| BioPerl | Perl tools for computational molecular biology |
| GROMACS | Versatile package to perform molecular dynamics |
| IGV | High-performance visualization genome browser tool |
| GATK | Genomic analysis toolkit focused on variant discovery |
| BioJava | Provides Java tools for processing biological data |
| InterMine | Integrate biological data sources |
| bedtools | Powerful toolset for genome arithmetic |
| EMBOSS | The European Molecular Biology Open Software Suite |
| BLAST | Algorithm for comparing primary biological sequence information |
| Galaxy | Web-based platform for data-intensive computational research |
| minimap2 | Versatile sequence alignment program |
| Jalview | Multiple sequence alignment editing, visualisation and analysis |
| samtools | Manipulate next-generation sequencing data |
| BCFtools | Variant calling and manipulating files in the Variant Call Format |
| FastQC | Quality control tool for high throughput sequence data |
| SPAdes | Versatile toolkit for assembling and analysing sequencing data |
| GenomeTools | Collection of bioinformatics tools |
| AliView | Alignment viewer and editor |
| mothur | Analyze microbial communities |
| Bandage | Visualising de novo assembly graphs |
| cramino | BAM/CRAM quality evaluation |
| abPOA | Adaptive banded Partial Order Alignment |
| Taverna Workbench | For designing and executing bioinformatics workflows |
| geWorkbench | Software platform for integrated genomic data analysis |
| Bioclipse | Rich-client platform chemistry and biology workbench |
Read our verdict in the software roundup.
Explore our comprehensive directory of recommended free and open source software. Our carefully curated collection spans every major software category.This directory is part of our ongoing series of informative articles for Linux enthusiasts. It features hundreds of detailed reviews, along with open source alternatives to proprietary solutions from major corporations such as Google, Microsoft, Apple, Adobe, IBM, Cisco, Oracle, and Autodesk. You’ll also find interesting projects to try, hardware coverage, free programming books and tutorials, and much more. Discovered a useful open source Linux program that we haven’t covered yet? Let us know by completing this form. |


Please read our Comment Policy before commenting.