caption-tool / README.md
stevenlearns's picture
Update README.md
458c525 verified
|
Raw
History Blame Contribute Delete
3.73 kB
---
license: mit
tags:
- image-captioning
- dataset-tools
- lora
- pytorch
pipeline_tag: image-to-text
library_name: transformers
pretty_name: Caption Tool
---
# Caption Tool
A small, friendly tool that writes text descriptions (captions) for every image
in a folder, all on your own computer.
- **No API key. No account. No internet** (after the first-time model download).
- **Private:** your images never leave your machine.
- **Lightweight:** runs on a GPU with as little as **2 GB of VRAM**.
- **Whole folders at once:** point it at a directory and it captions everything.
Made by [oohfixer.com](https://oohfixer.com).
---
## What it does
For every image it finds, Caption Tool writes a matching text file next to it:
```
vacation.jpg → vacation.txt
cat.png → cat.txt
```
That's it. The caption sits in the `.txt` file, ready to use for training
image models, organizing photos, or search.
---
## Install
You need **Python 3.13 or older** (3.14 doesn't work with this tool's
dependencies).
### One command (recommended)
Open a terminal and run:
```bash
curl -fsSL https://oohfixer.com/caption-tool/install.sh | bash
```
This downloads the tool, installs what it needs in its own isolated
environment, and adds a **`captions`** command you can run from anywhere. When it
finishes, open a new terminal (or run `source ~/.bashrc`).
Then just type:
```bash
captions
```
### Manual install
If you'd rather set it up by hand:
1. Download this folder and open a terminal in it.
2. Install the requirements:
```bash
pip install -r requirements.txt
```
3. Run it:
```bash
python caption.py
```
---
## Use it
Just run it:
```bash
python caption.py
```
It will ask you for a folder of images, show the settings, and caption them.
Sit back, a progress bar shows how far it's gotten.
### Want to skip the questions?
If you already know what you want, you can do it in one line:
```bash
python caption.py --path "/path/to/your/images"
```
Useful extras:
| Option | What it does |
| --- | --- |
| `--path "FOLDER"` | Caption this folder right away (no prompt). |
| `--force` | Re-caption images that already have a `.txt` file. |
| `--model florence2-base` | Pick a different model (see below). |
| `--max-new-tokens 512` | Allow longer captions. |
---
## Choosing a model
When you run it, you can pick which AI model writes the captions. Bigger = more
detailed, but uses more VRAM.
| Model | Size | Best for |
| --- | --- | --- |
| **florence2-large** (default) | ~1.5 GB | Slightly less accurate. Needs a normal GPU. |
| **florence2-base** | ~0.45 GB | **2 GB GPUs**, or when you want it fast. |
If your card only has 2 GB of VRAM, choose **florence2-base**.
---
## Prefix and suffix (optional)
If you're captioning images to train your own image model, you may want a
**trigger word** on every caption (for example, your subject's name). Set a
**prefix** and the tool adds it to the front of every caption automatically.
A **suffix** does the same at the end.
Leave them empty if you just want plain descriptions.
---
## Settings
Settings are stored in `caption_config.json` next to the script. You can edit
that file directly, or change things from the tool's menu. The defaults:
- Skip images that already have a caption (so re-running is safe).
- At least ~15 words per caption (it re-captions anything shorter).
- 256 token cap on caption length.
---
## Questions
- **Does it need to be online?** Only the very first time, to download the
model. After that, everything runs offline.
- **Where do my images go?** Nowhere. They're processed locally and the caption
is saved next to each image.
- **What image types?** jpg, png, webp, bmp, tiff.