Cohere Python SDK

Official Python client for the Cohere API, with typed access to chat, embed, rerank, and generate across AWS, Azure, GCP, and OCI.

SDK
PyPI
v7.2.0
404 stars
MIT License

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
75 /100 Good
Development Activity 60
Maintenance 68
Community 84
Maturity 60
Momentum 28

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
77 /100 Good
Architecture 80
Code Quality 75
Innovation 82
Learning Curve 70

The Cohere Python SDK is the official client library for accessing Cohere’s language models — chat, embed, rerank, classify, and generate — from Python code. It ships fully typed request and response models generated from Cohere’s API definition, so autocomplete and static type checkers understand every endpoint without hand-written stubs.

Beyond the default hosted platform, the SDK bundles dedicated clients for running Cohere models through AWS Bedrock and SageMaker, Azure, and Oracle Cloud Infrastructure, giving teams the same request/response shapes regardless of which cloud is actually serving the model. Sync and async usage are both first-class, streaming responses are supported for chat, and a local tokenizer integration lets callers count or split tokens without a network round trip.

What You Get

  • A synchronous Client and an AsyncClient sharing the same typed method signatures for chat, embed, rerank, classify, and generate
  • A ClientV2 targeting Cohere’s current chat API, including streaming via chat_stream
  • Platform-specific clients (BedrockClient, SagemakerClient, AzureClient, OciClient, and their async/v2 variants) for running the same calls against a partner cloud’s hosted model
  • Fully typed request and response models (pydantic-based) covering every documented API field, generated directly from Cohere’s API definition
  • A typed exception hierarchy (BadRequestError, UnauthorizedError, TooManyRequestsError, etc.) mapped from HTTP status codes for precise error handling
  • Local tokenizer utilities built on Hugging Face tokenizers for counting and splitting text without calling the API

Common Use Cases

  • Building a chat or RAG application against Cohere’s chat/chat_stream endpoints with typed message and citation objects
  • Generating and comparing text embeddings for search, clustering, or retrieval pipelines via embed
  • Reranking a candidate document list for search relevance using rerank
  • Running Cohere models from inside an AWS, Azure, or OCI environment without switching client libraries
  • Batch-processing large embedding jobs asynchronously with the dataset and embed-job utilities

Under The Hood

Architecture The SDK is layered: raw_base_client.py handles raw HTTP request construction and response parsing, base_client.py (auto-generated by Fern from Cohere’s API definition) builds typed BaseCohere/AsyncBaseCohere methods on top of it, and the hand-maintained client.py/client_v2.py wrap those with backwards-compatible method names, deprecation shims, local tokenizer integration, and response caching via manually_maintained/cache.py. Platform-specific clients (bedrock_client.py, sagemaker_client.py, oci_client.py) subclass the same base to swap only the transport and auth layer while keeping method signatures identical. Request/response types live under types/ as generated pydantic models, and errors are mapped from HTTP status codes to a typed hierarchy in errors/. Changing the generated base client would ripple into every platform-specific subclass and the hand-maintained wrappers that depend on its method signatures.

Tech Stack Built for Python 3.10+, packaged with Poetry (poetry-core build backend). httpx is the core HTTP client for both sync and async requests; requests is also a runtime dependency for select code paths. pydantic and pydantic-core back all generated data models, with fastavro used for embed-job dataset formats and Hugging Face tokenizers for local tokenization. oci (Oracle Cloud SDK) and aiohttp/httpx-aiohttp are optional extras enabled via pip install 'cohere[oci]' or cohere[aiohttp]. Linting and static analysis run through ruff and mypy with the pydantic mypy plugin.

Code Quality The package ships a py.typed marker and is fully typed throughout, with mypy configured in CI (.github/workflows/ci.yml) alongside ruff for lint and format checks. The tests/ directory has 14 files covering the sync client, async client, Bedrock, OCI (including an OCI-specific mypy test), dataset utilities, embed streaming, and client-initialization edge cases such as environment-variable auth fallback; several tests exercise real API calls rather than mocks, which is typical for a generated SDK test suite but means some tests require network/credentials rather than running fully isolated. Generated files carry an explicit “auto-generated by Fern” header, keeping the boundary between generated and hand-maintained code (the manually_maintained/ directory) unambiguous for contributors.

API Design The public surface favors small, consistent entry points: constructing cohere.ClientV2() and calling .chat()/.chat_stream() requires no boilerplate beyond an API key, which itself defaults to the CO_API_KEY environment variable. The standout design choice is deployment parity — BedrockClient, SagemakerClient, AzureClient, and OciClient expose the same method names and argument shapes as the default hosted client, so application code calling co.chat() or co.embed() doesn’t need to change when the underlying inference target moves to a different cloud. The tradeoff is that advanced usage requires familiarity with the generated type hierarchy under types/, since request objects are typed pydantic models rather than plain dicts.

Used by 5 apps in this directory

Python
100%
Apache 2.0

Agno

AI Development · Automation · Devops

42,644

Build, run, and manage agent platforms with a full production stack — SDK, runtime, and control plane included.

View details
93
Repo Health
87
Technical
66
Dependency
Built with
Python 100%
Updated today
Python
88%
Apache 2.0

Apache Airflow

Data Engineering

47,133

Define, schedule, and monitor complex data workflows as Python code — with a powerful UI, 80+ provider integrations, and battle-tested scalability across thousands of production deployments.

View details
96
Repo Health
89
Technical
64
Dependency
Built with
Python 88%
Updated today
Rust
42%
Apache 2.0

LanceDB

AI Development · Databases

11,628

Open-source, embedded vector database built on the Lance columnar format for fast multimodal search across billions of vectors, backed by Y Combinator (W23).

View details
90
Repo Health
86
Technical
72
Dependency
Built with
Rust 42%
Python 26%
HTML 23%
Updated today
Python
45%
GPL 3.0

MaxKB

AI Development · Knowledge Management

22,931

Build enterprise-grade AI agents with RAG, workflows & multi-modal support

View details
92
Repo Health
68
Technical
66
Dependency
Built with
Python 45%
Vue 37%
TypeScript 17%
Updated today
Python
65%
Other

Onyx

AI Agents · AI Assistants · Knowledge Management

32,380

Self-hostable AI platform with agentic RAG, 50+ connectors, deep research, code execution, and support for every major LLM provider.

View details
92
Repo Health
83
Technical
68
Dependency
Built with
Python 65%
TypeScript 25%
Updated today

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers