PEFT

Hugging Face's library of state-of-the-art parameter-efficient fine-tuning methods for large models.

Library
PyPI
v0.21.0
21,730 stars
Apache License 2.0

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
90 /100 Excellent
Development Activity 96
Maintenance 96
Community 76
Maturity 52
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
84 /100 Excellent
Architecture 86
Code Quality 87
Innovation 88
Learning Curve 76

PEFT (Parameter-Efficient Fine-Tuning) lets you adapt large pretrained models — language models, diffusion models, and more — to new tasks by training only a small fraction of their parameters instead of the full model. Methods like LoRA, prefix tuning, prompt tuning, and IA3 wrap a base model with lightweight adapter layers, cutting GPU memory and storage requirements dramatically (checkpoints shrink from gigabytes to megabytes) while achieving accuracy comparable to full fine-tuning.

PEFT is built and maintained by Hugging Face and integrates directly with Transformers (model.add_adapter), Diffusers (LoRA adapters for Stable Diffusion), Accelerate (distributed training), and TRL (RLHF fine-tuning), so it slots into workflows teams are likely already using rather than requiring a separate training stack.

What You Get

  • Implementations of state-of-the-art PEFT methods (LoRA, prefix tuning, prompt tuning, P-tuning, IA3, and more) under src/peft/tuners
  • A simple get_peft_model() wrapping API that turns any compatible Transformers/Diffusers model into a PEFT-trainable model
  • Direct integration with Transformers (add_adapter/load_adapter/set_adapter), Diffusers, Accelerate, and TRL
  • Support for combining PEFT with quantization (e.g. QLoRA) to fine-tune large LLMs on consumer-grade GPUs
  • Small, portable adapter checkpoints (often single-digit MBs) instead of full multi-gigabyte model copies per task

Common Use Cases

  • Fine-tuning a multi-billion-parameter LLM on a single consumer or prosumer GPU using LoRA or QLoRA
  • Training and swapping multiple task-specific adapters against one shared frozen base model to save storage
  • Fine-tuning Stable Diffusion or other diffusion models with LoRA adapters for custom styles or subjects
  • Applying RLHF/DPO fine-tuning to LLMs via TRL with PEFT adapters to keep training memory manageable

Under The Hood

Architecture - PEFT’s core lives in src/peft/tuners/ (one module per method — LoRA, prefix tuning, prompt tuning, IA3, and others — each implementing how adapter parameters are injected into and merged with a base model’s layers), src/peft/utils/ (shared helpers for parameter freezing, saving/loading adapters, and config handling), and src/peft/optimizers/ (PEFT-aware optimizer utilities). The PeftModel/get_peft_model() entry points wrap an arbitrary Transformers or Diffusers model, replacing target layers with adapter-augmented equivalents while keeping the base model’s weights frozen.

Tech Stack - Nearly pure Python (99.6% of the codebase) built directly on PyTorch, with tight integration points into transformers, diffusers, accelerate, and bitsandbytes (for quantized fine-tuning) rather than reimplementing model internals. A small amount of CUDA/C++ exists for specific low-level ops.

Code Quality - 63 test files exercise the tuners, config serialization, and integration points; the project has a fast, steady release cadence (32 releases, ~44 commits/month) reflecting active maintenance by a large contributor base (325 contributors) led by the Hugging Face team.

API Design - The core workflow is deliberately small: define a LoraConfig (or equivalent), call get_peft_model(model, peft_config), then train and save_pretrained() as usual — familiar to anyone who already uses Transformers’ Trainer. model.print_trainable_parameters() gives an immediate, concrete sense of how much compute is being saved, which lowers the barrier to trusting the technique on a first try.

Used by 11 apps in this directory

Python
59%
Apache 2.0

argilla

AI Development · Data Engineering

5,125

Collaborate on high-quality AI training data with a self-hosted annotation platform built for LLMs, NLP, and multimodal models.

View details
65
Repo Health
81
Technical
61
Dependency
Built with
Python 59%
Jupyter Notebook 21%
Updated 1 weeks ago
Python
92%
Apache 2.0

ART

AI Development

10,779

Give your LLM agents on-the-job training—ART lets you apply GRPO reinforcement learning to any multi-step agentic workflow with minimal code changes.

View details
85
Repo Health
82
Technical
73
Dependency
Built with
Python 92%
Updated 5 days ago
Python
100%
Apache 2.0

ClearML

Automation · Devops

6,892

Auto-magical MLOps platform that tracks experiments, versions data, orchestrates pipelines, and serves models with just two lines of code.

View details
94
Repo Health
79
Technical
68
Dependency
Built with
Python 100%
Updated 1 weeks ago
Python
83%
MIT

Gemma Multimodal Fine-Tuner

AI Development

1,503

An Apple-Silicon-native LoRA fine-tuning tool for Gemma on text, image, and audio data — with a wizard CLI, live browser-based training visualizer, and streaming from GCS/BigQuery for datasets too large for local disk.

View details
65
Repo Health
68
Technical
72
Dependency
Built with
Python 83%
Updated 6 days ago
C++
52%
MIT

GPT4All

AI Assistants · AI Development

77,382

Run large language models privately on your laptop — no GPU, no cloud, no data leaving your device.

View details
53
Repo Health
80
Technical
74
Dependency
Built with
C++ 52%
QML 30%
Updated 1 years ago
TypeScript
56%
Apache 2.0

LTX-Desktop

AI Design Tools · Video Editors

2,028

An open-source Electron app that runs LTX-2 text-to-video, image-to-video, and video editing models locally on your GPU, or via a cloud API when your hardware can't keep up.

View details
72
Repo Health
82
Technical
73
Dependency
Built with
TypeScript 56%
Python 41%
Updated 1 months ago
Python
62%
Apache 2.0

marimo

Data Engineering · Developer Tools

22,918

A reactive Python notebook that eliminates hidden state, runs reproducibly, and deploys as a web app or script — stored as pure Python, built for the AI era.

View details
89
Repo Health
91
Technical
65
Dependency
Built with
Python 62%
TypeScript 37%
Updated 5 days ago
Go
93%
MIT

NornicDB

AI Development · AI Memory · Databases

887

A single graph+vector+temporal database for AI workloads — Neo4j-compatible, sub-millisecond hybrid search, and built-in memory decay.

View details
80
Repo Health
82
Technical
71
Dependency
Built with
Go 93%
Updated 4 days ago
Rust
63%
MIT

PostgresML

AI Development · Databases

6,825

Run ML training and LLM inference natively inside PostgreSQL with GPU acceleration — no data movement required.

View details
51
Repo Health
76
Technical
64
Dependency
Built with
Rust 63%
JavaScript 11%
Updated 1 years ago

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers