All 22 Dependencies

Every package Gemma Multimodal Fine-Tuner depends on, ranked by repo health score.

Gemma Multimodal Fine-Tuner fills a specific gap: most LoRA fine-tuning tools for Gemma (MLX-LM, Unsloth, axolotl) handle text-only CSV well, but support image+text or audio+text fine-tuning unevenly or not at all, and typically assume an NVIDIA GPU. This tool runs natively on Apple Silicon via MPS (no CUDA required) and supports text-only, image+text (captioning/VQA), and audio+text LoRA fine-tuning from local CSV data.

For datasets too large to fit on a local SSD, it can stream training data directly from Google Cloud Storage or BigQuery rather than requiring everything downloaded first. A wizard CLI walks through system checks, LoRA configuration, model selection, and dataset setup, and a real-time training visualizer — loss curve, attention heatmap, gradient signal strength, memory pressure, and token-by-token predictions — runs live in the browser during training with a single config flag, no TensorBoard or notebook setup required.

MIT licensed and open source, it's positioned specifically for Mac-based ML practitioners who want multimodal Gemma fine-tuning without provisioning cloud GPU infrastructure.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers

Search