rank_bm25

A pure-Python implementation of the BM25 family of ranking algorithms for document search

Library
PyPI
v0.2.2
1,397 stars
Apache License 2.0

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
43 /100 Fair
Development Activity 8
Maintenance 20
Community 44
Maturity 60
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
68 /100 Good
Architecture 65
Code Quality 62
Innovation 55
Learning Curve 88

rank_bm25 implements several variants of the BM25 ranking function — Okapi BM25, BM25L, BM25+, and BM25-Adpt — used to score how relevant a document is to a search query based on term frequency and document length, without requiring a full search-engine stack like Elasticsearch or Lucene. Each variant is exposed as a class that indexes a tokenized corpus and returns relevance scores or top-N matches for a query.

Because it’s a small, dependency-light (NumPy-only) pure-Python library rather than a hosted search service, rank_bm25 is frequently used as a lightweight lexical-retrieval baseline or as the sparse-retrieval half of hybrid search pipelines that combine BM25 with dense vector/embedding search — a common pattern in retrieval-augmented generation (RAG) systems.

What You Get

  • BM25Okapi, BM25L, BM25Plus, and BM25Adpt classes implementing different BM25 scoring variants
  • get_scores() to score every document in the corpus against a tokenized query
  • get_top_n() to retrieve the top-N most relevant documents directly
  • A simple, transparent scoring model based on term frequency and inverse document frequency with length normalization
  • Minimal dependencies (NumPy only), making it easy to drop into any Python pipeline

Common Use Cases

  • Adding a lightweight keyword/lexical search baseline to a Python application without deploying Elasticsearch or a vector database
  • Implementing the sparse-retrieval half of a hybrid search pipeline alongside dense embedding-based retrieval in a RAG system
  • Re-ranking a small candidate document set by lexical relevance before or after a semantic-similarity pass
  • Teaching or prototyping information-retrieval concepts where a transparent, inspectable scoring formula matters more than production scale

Under The Hood

Architecture - The entire library lives in a single rank_bm25.py module: a BM25 base class handles corpus indexing (term frequencies, document lengths, average document length, inverse document frequency computation), and each variant (BM25Okapi, BM25L, BM25Plus, BM25Adpt) subclasses it to override only the scoring formula, keeping the differences between BM25 variants explicit and easy to compare.

Tech Stack - Pure Python with a single runtime dependency, NumPy, used for vectorized score computation across the corpus; there is no compiled extension, external index, or service dependency, so the whole library installs and runs anywhere NumPy does.

Code Quality - Test coverage in tests/test_loading.py is comparatively thin, focused mainly on corpus loading/scoring smoke tests rather than exhaustive coverage of every BM25 variant’s edge cases; the small single-file scope makes the implementation easy to read and verify manually against the published BM25 formulas, but the project has seen infrequent maintenance since its last 0.2.2 release.

API Design - Usage is a two-step pattern: construct BM25Okapi(tokenized_corpus) once to index the corpus, then call .get_scores(tokenized_query) or .get_top_n(tokenized_query, corpus, n) as many times as needed — a minimal, easy-to-learn API, though callers are responsible for their own tokenization since the library does no text preprocessing itself.

Used by 5 apps in this directory

Python
99%
MIT

Agent Lightning

AI Development

18,515

A Microsoft-built training framework that optimizes AI agents with reinforcement learning, automatic prompt optimization, or supervised fine-tuning — with near-zero code changes to your existing agent, in any framework.

View details
85
Repo Health
68
Technical
69
Dependency
Built with
Python 99%
Updated 4 days ago
Python
66%
Other

AutoGPT

AI Assistants · Automation · Productivity

187,596

Build, deploy, and run autonomous AI agents that automate complex multi-step workflows using a visual block-based graph editor.

View details
93
Repo Health
78
Technical
66
Dependency
Built with
Python 66%
TypeScript 33%
Updated 4 days ago
Python
79%
Apache 2.0

changedetection.io

Monitoring

34,605

Self-hosted website change detection with AI-powered smart alerts, browser automation, price tracking, and 85+ notification channels.

View details
91
Repo Health
80
Technical
68
Dependency
Built with
Python 79%
Updated 1 weeks ago
Python
86%
Apache 2.0

knowhere

AI Development · AI Memory · Developer Tools

3,541

Transform messy, unstructured documents into persistent, navigable memory that AI agents can actually use.

View details
82
Repo Health
75
Technical
66
Dependency
Built with
Python 86%
HTML 14%
Updated 1 weeks ago
Python
37%
Other

Open WebUI

AI Agents · AI Assistants

153,390

The extensible, privacy-first AI platform that runs Ollama, OpenAI, and any LLM backend behind a polished, feature-packed web interface.

View details
91
Repo Health
75
Technical
66
Dependency
Built with
Python 37%
Svelte 34%
JavaScript 21%
Updated 4 days ago

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers