Data Science & Numerical Computing Packages
Numerical computing, dataframes, and scientific-analysis libraries used across research and production data pipelines (NumPy, pandas, SciPy).
Packages in Data Science & Numerical Computing
Kind
Sort:Used by Most Apps
NumPy
The fundamental N-dimensional array library powering Python's scientific computing stack.
pandas
The Python DataFrame library for fast, flexible, and expressive data analysis and manipulation.
PyTorch
A Python-first tensor library with GPU acceleration and a dynamic, define-by-run autograd engine for building and training deep neural networks.
scikit-learn
Simple and efficient tools for machine learning and data analysis in Python.
PyArrow
Python bindings for Apache Arrow's columnar in-memory format, giving pandas, NumPy, and data pipelines a fast, zero-copy way to move and process tabular data.
Matplotlib
Python's foundational plotting library for static, animated, and interactive charts
SciPy
Fundamental algorithms for scientific computing, optimization, and statistics in Python.
networkx
Python library for creating, manipulating, and analyzing complex networks and graphs.
Arrow
The official Rust implementation of the Apache Arrow columnar in-memory data format.
Polars
Blazingly fast multi-threaded DataFrame library built in Rust with Python, Node.js, and R bindings
Apache Arrow JS
The official JavaScript/TypeScript implementation of Apache Arrow's columnar in-memory format for fast, zero-copy data interchange.
ipykernel
The IPython kernel that lets Jupyter notebooks, JupyterLab, and other frontends run and interact with Python code.
Streamlit
Turn a Python script into a shareable data app in minutes, with no frontend code
Seaborn
A Python statistical data visualization library built on matplotlib, for attractive charts with minimal code
SymPy
Full-featured computer algebra system (CAS) written entirely in pure Python
librosa
A Python library for audio and music signal analysis, providing the core DSP building blocks for music information retrieval (MIR) systems.
PySpark
Python API for Apache Spark, the unified engine for large-scale distributed data processing
yfinance
Pythonic access to Yahoo Finance market data, returned as ready-to-use pandas DataFrames.
Apache DataFusion
The extensible, Arrow-native SQL and DataFrame query engine for building fast analytical systems in Rust.
LightGBM
A fast, distributed, high-performance gradient boosting framework based on decision tree algorithms, used for ranking, classification, and other machine learning tasks.
statsmodels
Statistical modeling and econometrics in Python, with rigorous estimation and inference.
deltalake
Native Rust-powered Python binding for Delta Lake, giving pandas, Polars, and Arrow workflows ACID table reads and writes without a JVM.
scikit-image
A collection of algorithms for image processing in Python, built on NumPy arrays and interoperable with the scientific Python ecosystem.
GeoPandas
Add support for geographic vector data to pandas dataframes