1.3.0 Image preparation for AI creators

Prepare the images.
Then train.

Caption with Claude, OpenAI or local models. Upscale, tag and organize images for LoRA and other AI training workflows, all from one desktop app.

Free & open source / Windows 10 & 11, 64-bit

Local captioning · BLIP base

TrainKit turns chelsea.png into chelsea.txt with the caption: a photograph of a cat sitting on a bed.

Give Claude the context.
Get captions for your dataset.

Read the release notes ↗

A caption for a character LoRA needs different details than one for a product collection. Tell Claude what to describe, choose an image or a folder, and save a text caption for each image.

Connect your Anthropic API key in TrainKit’s API tab, then select Claude in Caption. You can also use OpenAI or run a local vision model.

See the API workspace
CAPTION INSTRUCTIONExample

Caption this image for a character LoRA. Describe the pose, clothing, expression, camera angle and background. Return a concise natural-language caption.

Your instruction is applied to each image in the batch.
  1. 01

    Choose your images

    One image or a folder, with your caption instruction.

  2. 02

    Claude reads each image

    TrainKit sends a resized image and your instruction directly to Anthropic.

  3. 03

    Keep the captions

    UTF-8 .txt files are saved in your chosen output folder.

Cloud captioning uses your provider’s API credits. Local captioning, upscaling, tagging and renaming run on your computer. How data is handled ↗

TrainKit 1.3.0 · API workspace
TrainKit 1.3.0 API tab with separate Anthropic Claude and OpenAI key fields, model settings and key controls
TrainKit v1.3.0 · Windows appOpen full resolution ↗

From an image to useful output.

Real inputs. Real runs.
Explore the results below.

Text that belongs with the image.

Use Claude, OpenAI or a local vision model to caption an image or a folder. TrainKit writes a UTF-8 text file for each image. This sample shows a local run with BLIP base.

chelsea.txt

a photograph of a cat sitting on a bed

BLIP base · prompt: “a photograph of”Download the sidecar ↓
See the captioning workspace
TrainKit 1.2.1 captioning interface, with BLIP loaded, the cat image selected and one file successfully processed
TrainKit v1.2.1 · Local captioning exampleOpen full resolution ↗

Sample photo: Chelsea by Stefan van der Walt, CC0 via scikit-image. Batch preview photo: NASA, public domain. These local processing examples were recorded with TrainKit v1.2.1.

Work through the whole folder.

Batch tools with progress
you can come back to.

Choose where captions run

Use local vision models or connect Claude and OpenAI. Cloud requests go straight to the provider; your output files are saved locally.

Resume a stopped run

Cancel a run and resume it using a saved manifest. Completed and skipped files keep their progress.

Handle existing files

Choose how output collisions are handled, keep track of each file’s status and preserve your original images.

Free or Pro.

Pick the tools that fit your workflow.

Choose the Pro billing period

Free

Available · 1.3.0

$0 to download

The open-source desktop app. Run local models or connect your own cloud API key.

  • Local captioning, upscaling, tagging and renaming
  • Claude and OpenAI captioning with your own API key
  • Batch processing and resumable runs
  • Full source code under the MIT license

Cloud usage is billed by your API provider.

Download Free

Pro

$25 / month

$240 / year · $20 per month

A customized TrainKit app with account login and hosted image processing.

  • Everything in Free
  • Cloud captioning through your TrainKit account
  • Cloud image upscaling
  • Access to more hosted models
  • 1,000 processing credits each month
Choose Pro

Put TrainKit to work.

Version 1.3.0 · Windows 10 / 11, 64-bit · Free · MIT licensed

Download 1.3.0
  1. Install the essentials

    Install uv and the Visual C++ x64 runtime.

  2. Extract and launch

    Unzip the release into a writable folder. Open TrainKit.exe and let setup finish.

  3. Choose your caption provider

    Connect your Claude or OpenAI API key, or select a local model. Choose your images and output folder, then start a run.

First-launch requirements

First-launch setup needs internet access and several gigabytes of free space. Local model files are selected separately. An NVIDIA GPU is recommended for local captioning and Spandrel; NCNN supports CPU and compatible Vulkan GPUs. Cloud captioning needs an API account with credits and does not require a local GPU.

Read the installation guide ↗