Write captions

Chapter 3

Write captions

One sentence per photo: what to write, what to leave out, and tools for many at once.

How captions work

One sentence per photo, starting with your trigger word.

Describe what should change from picture to picture. Leave out what the LoRA should learn.

The Dataset page with the Vell Trench example: 28 photos, three problems left in on purpose.
  1. Select a photo and type its caption in the box on the right. It saves when you leave the box or press ⌘S.
  2. Start every caption with the trigger word; it is highlighted as you type.
  3. Keep it short. The word and token count above the box turns yellow when a caption gets long.

Build a caption

Pose

Background

Light

Angle

vell_trench, standing, plain grey wall, soft window light, front view

Write what changes from photo to photo: the pose, the place, the light, the angle.

Leave out what the LoRA should learn: camel colourdouble-breastedbelted waiststorm flap

Change a word in every caption

Fix one phrase across all your captions at once.

The Dataset page with the Captions, Tag balance and Console panels open at the bottom.
  1. Click Find/Replace on the Dataset toolbar.
  2. Type the words to find and what they should become.
  3. Read the preview: every caption that changes is listed.
  4. Click Apply. ⌘Z undoes the whole batch.

Trigger word tools

Keep your trigger word spelt the same everywhere.

The Dataset page with photo 09 selected: its caption spells the trigger word vel_trench.
  1. Click Trigger word problems on the left to see captions that miss the word or spell it another way.
  2. Click Find/Replace and open the Trigger tab.
  3. Choose Rename from…, type the old spelling (vel_trench) and check that the trigger word field says vell_trench.
  4. Read the preview: each caption before and after. Only a whole comma-separated item that matches the old spelling (capitals or not) changes. If a caption then has the trigger word twice, only the first is kept.
  5. Click Apply. ⌘Z undoes the whole batch.

More on the Trigger tab

  • Add where missing puts the trigger word at the start of every caption that has none. It does not fix a misspelt one.
  • Move to the start moves the word to the front of each caption; Remove takes it out.

With Find & replace

Find & replace works for any text. On its tab, find the misspelt word (vel_trench), replace it with your trigger word (vell_trench) and tick Whole words, then check the preview and click Apply.

Try it

  • vell, trench coat, studio light, plain background
  • Vell, trench coat on a rainy street at night
  • vel, trench coat, close-up of the collar

1 / 3 captions use it exactly · 2 spelt another way

Tags and balance

See which words take over your captions.

The Dataset page with the Captions, Tag balance and Console panels open at the bottom.
  1. Open Tags under the grid. The bar shows what share of captions uses each word.
  2. Click a word to see the photos that use it.
  3. Right-click to rename or delete a word, or open All… to merge several into one.
  4. Drag words to reorder them; the trigger word stays first.

If one background or pose is in most captions, the LoRA will learn it too. Rewrite those captions or shoot more variety.

Write many captions at once

Add the same words to a group of photos.

The Dataset page with the Mira K example: 22 photos of one character.
  1. Select the photos.
  2. Click Find/Replace to open the batch caption tools, and type what to add or change.
  3. Check the preview and apply. ⌘Z undoes it.

Coming later

  • AI caption drafts written on your Mac.
  • Importing captions from another AI.
  • Reviewing new captions one by one.

AI models on your Mac

Small AI models that draft captions and tags and cut masks on your Mac, without uploading your photos.

Coming later

None of this is in the first beta. It is here so that you can plan the disk space and check the licences.

Which models

Skye Desk offers a small first set: about 3.6 GB in all, with the Python they run in (about 1.5 GB).

ModelDownload sizeLicenceCommercial use
Florence-2-large-PromptGen v2.0Caption drafts in sentences or tags; finds the clothes, face or hair for a maskabout 1.6 GBMITAllowed
WD vit tagger v3Comma tags for SDXL projectsabout 0.4 GBApache-2.0Allowed
SAM 2.1 tinyCuts the masksabout 0.2 GBApache-2.0Allowed

Larger models are listed too, for when you have the space:

ModelDownload sizeLicenceCommercial use
JoyCaption Beta OneCaption draftsabout 16.5 GBLlama 3.1 Community LicenseWith conditions
Qwen2.5-VL-7B-InstructCaption draftsabout 16.6 GBApache-2.0Allowed
Qwen3-VL-4B-InstructCaption draftsabout 8.9 GBApache-2.0Allowed
Florence-2-largeFinds the parts for masksabout 1.5 GBMITAllowed
WD eva02-large tagger v3Tagsabout 1.2 GBApache-2.0Allowed
BiRefNetMasks of the whole subjectabout 0.9 GBMITAllowed
SAM 2.1 base+Cuts the masksabout 0.3 GBApache-2.0Allowed

JoyCaption follows the Llama 3.1 Community License, which adds conditions – credit, naming and an acceptable-use policy. Read the licence text before you use it. Models whose licence does not allow commercial use, such as RMBG-2.0, Qwen2.5-VL-3B and the clothes and face parsing models, are not offered.

Download and remove

  • Nothing downloads until you confirm. First Skye Desk shows each model’s name, size and licence, and how much free space will be left.
  • Models come from Hugging Face. A download can be stopped and resumed, and every file is checked before it is used.
  • They live in their own folder, ~/Skye Desk/models, outside your workspace and with their own Python: nothing else on your Mac changes.
  • To free the space, quit Skye Desk and delete the model’s folder there.

What needs them

  • Auto caption on the Dataset toolbar writes drafts for the photos you choose. They are a preview: nothing reaches your captions until you click Apply, and ⌘Z undoes the whole batch.
  • Masks cuts out the clothes, the face or the hair, so that training can focus on them. Masks of the whole subject need BiRefNet from the larger set.
  • Without the models everything else works as before, and you write the captions yourself.

Captions for each kind

What to write for clothing, characters, hair, products and styles.

Clothing

my_coat, a trench coat on a dress form, side view, plain grey background

Describe: the pose, the person or dress form, the background, the light, the camera angle.

Leave out: the fabric, colour, cut and details you want the coat to keep – the LoRA should learn those.

Characters

my_char, sitting on a bench, three-quarter view, soft evening light, city street

Describe: what the character is doing, the pose, the clothes if they should change, the place and the light.

Leave out: the face, the hair, the colours and markings that make the character who they are.

Hair

my_hair, woman seen from the side, indoor window light, plain wall

Describe: who wears it, the angle, the light, the background, the clothes.

Leave out: the cut, the length, the colour and the texture of the hair itself.

Products

my_bottle, on a stone plinth, top view, morning light, soft shadow

Describe: where it stands, the angle, the light, the props and the background.

Leave out: the shape, the label, the colours and the logo – the parts that must stay the same.

Styles

my_style, a lighthouse on a cliff, wide view

Describe: the subject of each picture – it should be different in every one.

Leave out: how it looks: the brush, the palette, the line – that is what the LoRA learns.

Next chapterTrain and chooseSet up a run, check it before it starts, and keep the best version.

Image preview