Auto-label any image dataset with SAM 3 using plain-English prompts. A full Python pipeline covering Hugging Face, dataset prep, masks, and COCO/YOLO export.
How to auto-label your dataset using SAM 3
Annotation is where most computer vision projects quietly die. You have 40,000 images of warehouse footage, a clear idea of what you want to detect, and a quote from a labelling vendor that costs more than the rest of the project combined. So the dataset sits there.
SAM 3 changes the shape of that problem. Give it a short noun phrase such as
forklift, shipping pallet, or person in hi-vis vest, and it will find and
segment every instance of that concept in an image, with no training, no
bounding-box seeds, and no class list it was pretrained on. That is good enough
to generate training labels directly.
This post is the full pipeline: pull a dataset off the Hugging Face Hub, prepare it properly (the step that actually decides your label quality), run SAM 3 with text prompts, calibrate the thresholds, and export to COCO or YOLO. Every snippet below is from a working repo, and at the end there’s a version that skips the GPU entirely.
What SAM 3 actually does differently
If your mental model of Segment Anything is “click a point, get a mask,” SAM 3 is a different tool. Meta introduced it in SAM 3: Segment Anything with Concepts with a task they call Promptable Concept Segmentation (PCS): given a short noun phrase, return instance masks and identities for every object matching that concept.
The distinction matters for auto-labelling:
| SAM 2 | SAM 3 | |
|---|---|---|
| Prompt | Point, box | Point, box, text phrase |
| Output | One object per prompt | Every matching instance |
| Open vocabulary | Limited | Yes |
| Needs a seed click | Yes | No |
“Every matching instance” is the part that makes unattended labelling possible.
With SAM 2 you still needed a human (or another detector) to say where to
look. With SAM 3 the prompt forklift returns all six forklifts, each as a
separate mask with its own confidence score.
A few specifics worth knowing before you build on it:
- 848M parameters, a DETR-based detector and a SAM 2-style tracker sharing one vision encoder. A presence head decouples recognition (“is this concept in the image at all?”) from localisation, which is a large part of why it handles negative prompts well.
- Native resolution is 1008 px. The processor resizes to this. Feeding larger images buys you nothing; feeding smaller ones loses small objects.
- ~30 ms per image on an H200 with 100+ objects, per Meta’s announcement. Real-world end-to-end throughput including I/O is lower. Plan for a few images per second per GPU.
- 48.8 zero-shot mask AP on LVIS, which is the number that justifies using the raw output as labels instead of as suggestions.
- The weights are gated and the licence is custom. SAM 3 ships under Meta’s SAM License, not MIT or Apache. Read it before you ship.
The pipeline
prepare_dataset.py → label_local.py ─┐
label_api.py ─┴→ calibrate.py → export.py → review.py
manifest.jsonl labels.jsonl thresholds COCO/YOLO contact sheets
Six stages, one shared intermediate format. The two labelling backends (your
own GPU or a hosted API) are interchangeable because they emit identical
labels.jsonl files. Everything downstream is written once.
Prerequisites
python -m venv .venv && source .venv/bin/activate
pip install torch torchvision transformers accelerate datasets \
pillow numpy pycocotools opencv-python-headless tqdm requests
For the local backend you need a GPU with 16 GB of VRAM or more and access
to the gated facebook/sam3 repo. Accept the terms on the
model page, then:
huggingface-cli login
Approval is usually quick but is not instant. If you’re on a deadline, start that request before you write any code.
Step 1: Pull a dataset from Hugging Face
Any dataset with an image column works. The Hub’s datasets library gives you
PIL images directly:
from datasets import load_dataset
dataset = load_dataset("keremberke/forklift-object-detection", split="train")
print(dataset.features)
# {'image': Image(mode=None, decode=True), 'image_id': Value('int64'), ...}
For anything large, stream it rather than downloading the whole archive:
dataset = load_dataset("your/dataset", split="train", streaming=True)
Column names are not standardised across the Hub, so resolve the image column by type rather than by guessing at a name:
def find_image_column(dataset) -> str:
"""Return the name of the column holding PIL images."""
features = getattr(dataset, "features", None)
if features:
for name, feature in features.items():
if type(feature).__name__ == "Image":
return name
# Streaming datasets may not expose features; fall back to a sample.
sample = next(iter(dataset))
for name, value in sample.items():
if isinstance(value, Image.Image):
return name
raise SystemExit("Could not find an image column.")
Step 2: Prepare the dataset (this is the step people skip)
Here’s the uncomfortable truth about auto-labelling: the model is rarely your bottleneck. Your input pipeline is. SAM 3 will happily produce confident, well-formed, completely useless masks on a badly prepared dataset, and because nothing crashes you won’t find out until you’ve trained on them.
Five things to fix, in order of how much damage they do.
2.1 Resolution: match SAM 3’s native 1008 px
SAM 3 operates at 1008 px. The HF docs are explicit that custom resolutions may degrade accuracy. So:
- Downscale the long edge to 1008. Anything larger is decoded, resized, and thrown away. You’re paying JPEG decode cost for pixels the model never sees.
- Never upscale. A 320 px thumbnail upscaled to 1008 px has no more information than it did; it just takes longer to process and produces confidently mushy mask boundaries.
- If your objects are genuinely tiny (aerial imagery, PCB inspection, microscopy), don’t resize the whole frame. Tile it. A 4000 px image squashed to 1008 px turns a 40 px defect into a 10 px smudge that no segmentation model will recover. Tile into overlapping 1008 px crops, label each, and merge the masks back with the tile offset.
SAM3_NATIVE_SIZE = 1008
def normalise(image: Image.Image, target_long_edge: int) -> Image.Image:
"""EXIF-correct, force RGB, and clamp the long edge. Never upscales."""
image = ImageOps.exif_transpose(image) # honour camera rotation
if image.mode != "RGB":
image = image.convert("RGB") # greyscale/RGBA/palette -> RGB
long_edge = max(image.size)
if long_edge > target_long_edge:
scale = target_long_edge / long_edge
new_size = (round(image.width * scale), round(image.height * scale))
image = image.resize(new_size, Image.LANCZOS)
return image
2.2 Colour mode and EXIF rotation
Two silent corruptors:
- EXIF orientation. Phone photos are frequently stored rotated with a flag
telling the viewer to turn them. PIL does not apply that flag automatically.
Your masks come back correct for the stored orientation and wrong for the
one you see in a viewer.
ImageOps.exif_transpose()fixes it, and it must happen before anything else. - Channel count. Greyscale medical scans, RGBA screenshots with an alpha channel, and palettised PNGs all need converting to RGB. An unconverted alpha channel in particular produces bizarre, hard-to-debug results.
2.3 Duplicates
Public datasets and scraped folders are full of near-duplicates: the same frame at two compression levels, the same product shot with a different watermark. Duplicates inflate your image count, cost you inference budget, and leak between train and validation splits, so your model looks better than it is.
A difference hash catches these cheaply:
def dhash(image: Image.Image, size: int = 8) -> str:
"""Difference hash - cheap perceptual fingerprint for near-duplicate removal."""
small = image.convert("L").resize((size + 1, size), Image.LANCZOS)
pixels = np.asarray(small, dtype=np.int16)
bits = pixels[:, 1:] > pixels[:, :-1]
return f"{int(''.join('1' if b else '0' for b in bits.flatten()), 2):016x}"
Exact hash collisions catch re-encodes. For crops and watermarks, compare Hamming distance and treat anything within 4–6 bits as a duplicate.
2.4 Degenerate and corrupt images
Truncated JPEGs, 20×20 favicons that snuck into a scrape, all-black frames from a camera dropout. Filter on a minimum short side and wrap the decode in a try/except so a single corrupt file doesn’t kill a six-hour run:
for source_id, loader in stream:
try:
image = normalise(loader(), args.long_edge)
except Exception as exc: # truncated / unreadable file
stats["corrupt"] += 1
print(f"[skip] corrupt {source_id}: {exc}")
continue
if min(image.size) < args.min_side:
stats["too_small"] += 1
continue
2.5 Reserve a dev slice before you label anything
Set aside 50–100 images that you will label first and look at with your own eyes. This is what you’ll use to tune prompts and thresholds. Choosing a threshold by labelling all 40,000 images and squinting at the aggregate is how you burn a day and a GPU budget on a prompt that was wrong from the start.
The output of stage one is a manifest.jsonl that every later stage reads:
{
"image_id": 0,
"file_name": "000000.jpg",
"width": 1008,
"height": 756,
"source_id": "keremberke/forklift-object-detection#0",
"phash": "b09c9292532c0c04",
"split": "pool"
}
Running it on a messy folder shows exactly what it caught:
$ python prepare_dataset.py --local-dir raw --out data/forklift --dev-slice 50
[skip] corrupt raw/broken.jpg: cannot identify image file 'raw/broken.jpg'
[prepare] read=15 corrupt=1 too_small=1 duplicate=1 kept=12 (dev=3)
[prepare] wrote data/forklift/manifest.jsonl
Step 3: Turn your class list into prompts
SAM 3 takes short noun phrases, not instructions. This is the single highest-leverage thing in the pipeline and it takes ten minutes to get right on your dev slice.
| Do | Don’t |
|---|---|
forklift |
find all the forklifts in this image |
shipping pallet |
pallets (plural dilutes the concept) |
person in hi-vis vest |
worker (too abstract, poorly grounded) |
red handbag |
the red one (no referent) |
Rules that hold up in practice:
- Singular, concrete nouns.
car, notcars, notvehicles present. - Modifiers work and are the main tool for narrowing.
white bicycleandbicyclereturn meaningfully different sets. If a prompt over-fires, add an adjective before you touch the threshold. - One concept per prompt.
car or truckis not a concept. Run two prompts. - Test the negative case. Run each prompt against 10 images you know contain none of that class. A prompt that fires on empty images will poison your dataset with false positives at scale. SAM 3’s presence head is good at this, but domain-specific jargon can still ground badly.
- Prefer visual descriptions over domain terms.
person in hi-vis vestbeatsPPE compliant workerbecause the model was trained on how things look, not on your industry’s vocabulary.
Step 4: Run SAM 3
The minimal call is short:
import torch
from PIL import Image
from transformers import Sam3Model, Sam3Processor
model = Sam3Model.from_pretrained("facebook/sam3").to("cuda").eval()
processor = Sam3Processor.from_pretrained("facebook/sam3")
image = Image.open("frame.jpg").convert("RGB")
inputs = processor(images=image, text="forklift", return_tensors="pt").to(model.device)
with torch.inference_mode():
outputs = model(**inputs)
results = processor.post_process_instance_segmentation(
outputs,
threshold=0.5, # detection confidence
mask_threshold=0.5, # mask binarisation
target_sizes=inputs["original_sizes"].tolist(),
)[0]
print(f"Found {len(results['masks'])} objects")
# results["masks"] -> binary masks at original image size
# results["boxes"] -> xyxy, absolute pixels
# results["scores"] -> confidence per instance
The optimisation that matters at scale
Written naively, a multi-class run encodes each image once per prompt. With three classes over 40,000 images that’s 120,000 passes through an 848M-parameter vision backbone. Of those, 80,000 are recomputing an embedding you already had.
SAM 3 lets you compute the vision features once and condition on each text prompt separately:
def label_image(model, processor, image, prompts, threshold, mask_threshold,
min_area, autocast) -> list[dict]:
"""Run every prompt against one image, reusing the vision embedding."""
image_inputs = processor(images=image, return_tensors="pt").to(model.device)
target_sizes = image_inputs["original_sizes"].tolist()
with torch.inference_mode(), autocast():
vision_embeds = model.get_vision_features(
pixel_values=image_inputs.pixel_values
)
annotations = []
for prompt in prompts:
text_inputs = processor(text=prompt, return_tensors="pt").to(model.device)
with torch.inference_mode(), autocast():
outputs = model(vision_embeds=vision_embeds, **text_inputs)
results = processor.post_process_instance_segmentation(
outputs,
threshold=threshold,
mask_threshold=mask_threshold,
target_sizes=target_sizes,
)[0]
masks = results["masks"].cpu().numpy()
scores = results["scores"].cpu().numpy()
for mask, score in zip(masks, scores):
annotation = build_annotation(mask, prompt, score, min_area=min_area)
if annotation:
annotations.append(annotation)
return annotations
The backbone runs once; only the (cheap) text conditioning and mask decoding repeat. On a multi-class job this is close to an N-fold speedup.
The single-prompt case has a similar optimization: precompute
model.get_text_features() once and reuse it across every image:
text_inputs = processor(text="forklift", return_tensors="pt").to(model.device)
with torch.inference_mode():
text_embeds = model.get_text_features(**text_inputs)
# then per image:
outputs = model(
pixel_values=img_inputs.pixel_values,
text_embeds=text_embeds,
attention_mask=text_inputs.attention_mask, # must be passed alongside
)
Also wrap inference in bf16 autocast on CUDA. It roughly halves memory with no measurable mask quality cost:
def autocast():
if device.type == "cuda":
return torch.autocast("cuda", dtype=torch.bfloat16)
return nullcontext()
Step 5: Encode masks so they survive contact with a trainer
A raw boolean array per instance is unusable at scale. At 40,000 images × 8 instances × 1008 × 756 booleans is hundreds of gigabytes. Run-length encoding gets that to a few hundred megabytes with no loss:
from pycocotools import mask as mask_utils
def encode_mask(binary_mask: np.ndarray) -> dict:
"""Binary HxW array -> COCO RLE with a JSON-safe (str) counts field."""
rle = mask_utils.encode(np.asfortranarray(binary_mask.astype(np.uint8)))
rle["counts"] = rle["counts"].decode("ascii")
return {"size": [int(binary_mask.shape[0]), int(binary_mask.shape[1])],
"counts": rle["counts"]}
Two details that bite people:
np.asfortranarrayis required.pycocotoolsexpects column-major memory; pass a C-ordered array and you get a silently transposed mask.- The
countsfield comes back asbytes, whichjson.dumpsrefuses. Decode to ASCII on the way out, re-encode on the way in.
Every annotation lands in one line-delimited JSON record per image:
{
"image_id": 0,
"file_name": "000000.jpg",
"width": 1008,
"height": 756,
"backend": "local",
"annotations": [
{
"category": "forklift",
"score": 0.93,
"bbox": [412.0, 188.0, 260.0, 341.0],
"area": 18342,
"segmentation": { "size": [756, 1008], "counts": "..." }
}
]
}
Step 6: Calibrate the threshold without re-running the model
threshold is a filter over confidence scores. So if you label your dev slice
once at a deliberately low threshold (0.15), you can sweep every candidate
value offline with no GPU or re-inference:
python label_local.py --data data/forklift --split dev \
--prompts "forklift" --threshold 0.15 \
--out data/forklift/labels_dev.jsonl
python calibrate.py --labels data/forklift/labels_dev.jsonl
=== forklift ===
thresh dets per img empty med area% p10 score
0.20 14 1.17 0.0% 3.23% 0.267
0.50 9 0.75 25.0% 2.82% 0.557
0.80 5 0.42 58.3% 4.51% 0.894
How to read it:
per imgshould land near the true object count for your domain. If you know warehouse frames average two forklifts and you’re getting 0.4, your threshold is too high or your prompt is wrong.- A sharp jump in
detsas the threshold drops is the false-positive cliff. Stop just above it. med area%collapsing as you lower the threshold means the extra detections are specks. Raise--min-areainstead of the threshold. You want the real low-confidence objects, not the noise.- High
emptyat every threshold means the prompt is wrong. No threshold will save it. Rewrite the noun phrase and re-run the dev slice.
The mask threshold can’t be swept this way because it changes pixels, not scores. Leave it at 0.5 unless masks are visibly too tight (lower it) or bleeding into background (raise it).
Expect different thresholds per class. forklift might settle at 0.45 while
person in hi-vis vest needs 0.65. Store them per prompt.
Step 7: Export to COCO or YOLO
COCO is a direct mapping. RLE segmentation drops straight in:
coco["annotations"].append({
"id": annotation_id,
"image_id": record["image_id"],
"category_id": category_ids[annotation["category"]],
"segmentation": annotation["segmentation"], # COCO RLE
"bbox": annotation["bbox"],
"area": annotation["area"],
"iscrowd": 0,
})
YOLO segmentation wants normalised polygons instead, so contour the mask and simplify:
def mask_to_polygons(binary_mask, epsilon_ratio=0.002, min_points=3):
"""Largest external contours, simplified with Douglas-Peucker."""
contours, _ = cv2.findContours(
binary_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
polygons = []
for contour in contours:
perimeter = cv2.arcLength(contour, True)
simplified = cv2.approxPolyDP(contour, epsilon_ratio * perimeter, True)
if len(simplified) >= min_points:
polygons.append(simplified.reshape(-1, 2))
return polygons
Two things to get right:
epsilon_ratiois a real tradeoff. Too low and you write 400-point polygons that bloat your labels; too high and you sand the corners off. 0.002 of the perimeter is a good default.- Write an empty
.txtfor images with no detections. YOLO treats a missing label file as missing data, but an empty one as a confirmed negative, and negatives are how the model learns not to hallucinate.
python export.py --data data/forklift --format coco --out exports/coco
python export.py --data data/forklift --format yolo --out exports/yolo --min-score 0.45
Step 8: Look at the labels before you train on them
Auto-labelling fails quietly. Metrics won’t tell you that your pallet prompt
has been segmenting the wooden floor for 12,000 images. Your eyes will, in about
ten seconds.
python review.py --data data/forklift --sort-by score --n 24 --out review/
Sorting by lowest confidence first puts the failures on screen immediately instead of showing you 24 easy successes. Check three things:
- Is the concept right? Are the masks on the thing you asked for?
- Are instances separated? Two adjacent forklifts as one blob means you need a more specific prompt, or a human pass on crowded frames.
- Are boundaries tight? Consistent bleed into background is a
mask_thresholdproblem, not a prompt problem.
Common failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Masks on the right objects but boundaries are mush | Images upscaled to 1008 px | Don’t upscale; tile large images instead |
| Confident masks on empty images | Prompt grounds to something else | Add a modifier; test the prompt on known-negative images |
| Small objects entirely missed | Whole frame downscaled | Tile into overlapping 1008 px crops |
| Adjacent instances merged into one | Genuine PCS limitation on dense scenes | More specific prompt, or human pass on crowded frames |
| Masks correct but rotated/flipped | EXIF orientation not applied | ImageOps.exif_transpose() before anything else |
| Silently transposed masks in export | C-ordered array into pycocotools |
np.asfortranarray() |
| Great on public data, poor on yours | Domain shift (thermal, medical, aerial) | Re-test prompts on your dev slice; consider fine-tuning |
That last row deserves emphasis. SAM 3’s zero-shot numbers come from natural imagery. On thermal, radiological, or synthetic-aperture data it can degrade sharply. Always validate on 50 of your images before committing to a full run.
What this costs on your own GPU
The honest accounting, because the compute is not the expensive part.
Compute. SAM 3 needs 16 GB of VRAM to be comfortable. At a realistic 3–5 images/second end-to-end on an A100-class GPU, 10,000 images with three prompts (using the vision-embedding reuse above) is roughly 45–90 minutes. At typical on-demand rates that’s a few dollars.
Everything else. Gated repo approval. A 3.4 GB weight download. CUDA and driver version alignment. The first OOM at batch size 2 and the hour of tuning after it. Queue management so a corrupt file at image 9,000 doesn’t lose the run. Then keeping that environment working three months later when you want to label another 5,000 images.
For most teams that’s half a day to a day of engineering the first time. If you’re labelling millions of images on a recurring basis, absolutely pay it once the marginal compute cost is unbeatable, so you should self-host.
If you’re labelling ten thousand images once, to find out whether the idea even works, you’re paying a day of setup for ninety minutes of compute.
Skip the GPU: the same pipeline as an API call
This is the part where the pipeline stops needing a GPU at all. SegmentationAPI runs SAM 3 as a managed service: you upload, submit prompts, and get masks back. Same model, same text prompts, same quality. No weights, CUDA, gated repo, or operations burden.
The flow is three calls (full reference):
1. Upload. Request a presigned URL, PUT the bytes straight to storage:
presign = session.post(f"{API}/uploads/presign",
json={"contentType": "image/jpeg"}).json()
requests.put(presign["uploadUrl"], data=path.read_bytes(),
headers={"Content-Type": "image/jpeg"})
task_id = presign["taskId"]
2. Segment. Submit up to 100 images per job with your prompts. This is the same prompt list you tuned on your dev slice, unchanged:
job_id = session.post(f"{API}/jobs", json={
"type": "image",
"tasks": task_ids, # up to 100 per job
"prompts": ["forklift", "shipping pallet", "person in hi-vis vest"],
"threshold": 0.45,
"maskThreshold": 0.5,
}).json()["jobId"]
3. Retrieve. Poll until the job succeeds, request the results archive, and read the manifest:
self.session.post(f"{API}/jobs/{job_id}/download")
while True:
info = self.session.get(f"{API}/jobs/{job_id}/download").json()
if info["status"] == "ready":
break
time.sleep(info.get("retryAfterSeconds") or 3)
archive = zipfile.ZipFile(io.BytesIO(requests.get(info["downloadUrl"]).content))
manifest = json.loads(archive.read(
next(n for n in archive.namelist() if n.endswith("output_manifest.json"))))
Each item in the manifest carries per-instance masks with confidences:
{
"taskId": "0000_a1b2c3d4",
"masks": [
{ "maskIndex": 0, "confidence": 0.97, "url": "masks/0000_a1b2c3d4/0.png" },
{ "maskIndex": 1, "confidence": 0.91, "url": "masks/0000_a1b2c3d4/1.png" }
]
}
Decode those into the same build_annotation() call the local backend uses, and
every downstream stage is unchanged: same labels.jsonl, same
calibrate.py, same COCO and YOLO exports. In the repo this is label_api.py,
a drop-in swap for label_local.py:
export SEGMENTATIONAPI_KEY=sk-your-key
python label_api.py --data data/forklift \
--prompts "forklift" "shipping pallet" "person in hi-vis vest"
No torch in your dependency tree at all.
When to use which
Both are reasonable. The choice depends on volume and on what your team’s time is worth:
Use the API when you’re validating an idea, your dataset is in the thousands to low tens of thousands, your labelling is bursty, you don’t want a GPU environment to maintain, or you need labels this afternoon. At $0.02 per processed image, ten thousand images is $200, less than the day of engineering it replaces, and you get 30-day asset retention and automatic SAM 3 upgrades without a migration.
Self-host when you’re running six or seven figures of images on a recurring schedule, you already have GPU infrastructure and someone who owns it, or your data can’t leave your network. At that scale the marginal compute cost wins clearly, and the setup cost amortises to nothing.
One cost lever worth knowing either way: billing is per processed image, so a single job carrying all three prompts costs one token per image, not three. Only split into per-prompt jobs when you need each mask attributed to a specific class.
The nice property of building the pipeline this way is that the decision isn’t permanent. Prep, calibration, export, and review are backend-agnostic. Start on the API, and if you outgrow it, swap one file.
FAQ
Can SAM 3 auto-label any dataset? Any dataset where your targets can be described as short noun phrases and look roughly like things in natural imagery. It’s excellent on retail, warehouse, traffic, agriculture, and general object data. For specialist modalities such as thermal, radiology, or satellite SAR, validate on a small dev slice before you commit; zero-shot quality can drop sharply under domain shift.
Are auto-labels good enough to train on? For most applications, yes. SAM 3’s 48.8 zero-shot mask AP on LVIS is close enough to human annotation that models trained on its output perform well. The standard workflow is auto-label everything, then have a human review the low-confidence tail, typically 5–15% of instances. That’s an order of magnitude less human time than labelling from scratch.
Do I need a GPU? For the local backend, yes. You need at least 16 GB of VRAM. SAM 3 is an 848M-parameter model designed for GPU inference; CPU inference is technically possible and practically far too slow for a dataset. The API backend needs no GPU.
How do I label video instead of images?
SAM 3 includes a SAM 2-style tracker that maintains object identity across
frames, so one text prompt segments and tracks through an entire clip. In
transformers that’s Sam3VideoModel with processor.add_text_prompt(); on
the API it’s "type": "video" with fps or numFrames sampling.
What about SAM 3.1?
Meta released improved SAM 3.1 checkpoints with Object Multiplex, a shared-memory
approach to multi-object tracking that’s dramatically faster on video. At the
time of writing they require the model code from the facebook/sam3 GitHub repo
rather than the transformers integration, and the image-labelling API shown
here is unchanged.
Can I use SAM 3 commercially? SAM 3 ships under Meta’s custom SAM License, not MIT or Apache. It permits building products, but attaches obligations to redistribution of the weights and derivatives, and forbids certain uses outright. Read the licence before shipping shipping. These terms govern the model, not the labels you generate with it.
How much does auto-labelling cost? Self-hosted, the marginal cost is GPU time, often dollars per ten thousand images, after a day of setup. Hosted, it’s $0.02 per processed image with no setup. The crossover is roughly wherever a day of your engineering time stops being worth more than the API bill.
Wrapping up
The interesting shift here isn’t that segmentation got better. It’s that the interface to segmentation became a text box. Once “find every instance of this concept” is a string, dataset creation stops being a procurement problem and becomes a scripting problem, and the bottleneck moves from annotation budget to how carefully you prepared your images and phrased your prompts.
Which is a much better problem to have.
Get the full pipeline from the examples repo, or grab an API key and label your first dataset without provisioning anything.