AI Tagging Cost Guide. Library Sizes 1k → 1M
What it actually costs to product-tag a UGC library with a vision model, from 1k to 1M assets: the token maths, the worked numbers at current list prices, and the review labour that usually costs more than the model.
- 12 min read
- For: cto, ecommerce leader
- Cited in AI search
- Star eligibility
- Tracked vs holdout
AI Tagging Cost Guide. Library Sizes 1k → 1M
What you’ll learn
- How two-stage tagging works: a cheap shortlist step, then one vision call per asset
- Per-asset token maths you can check against your own images and catalogue
- Worked cost projections at 1k, 10k, 100k and 1M assets, at list prices checked on 24 September 2026
- Why the human review queue, not the model bill, is usually the bigger cost, and how to size it
Chapter previews
- Chapter 01
Two-stage tagging explained
Stage one narrows a large catalogue to a shortlist cheaply. Stage two sends the image and that shortlist to a vision model once. How Idukki runs it, and where the confidence thresholds sit.
- Chapter 02
Cost at scale
A worksheet you fill with your own library size, image resolution and catalogue size, with worked numbers at current Gemini list prices from 1k to 1M assets.
- Chapter 03
When to keep humans in the loop
Route by confidence band, size the review queue from your own pilot, and measure where vision disagrees with a human on your catalogue rather than borrowing someone else's error rate.
Inside the playbook
In this article
Product tagging is the step that turns a customer photo into something a shopper can buy from: the post is linked to the SKUs visible in it, so the gallery tile, the hotspot and the add-to-cart all point at the right product. Doing it by hand is fine at fifty posts a month and hopeless at five thousand. Vision models have made automatic tagging cheap enough to run on every asset, but "cheap" hides a few decisions that move the bill by an order of magnitude. This worksheet walks through them with real list prices, so a CTO or Head of Ecommerce can size the job before committing to build or buy.
A note on the numbers. Every price below comes from Google's public Gemini API pricing page, checked on 24 September 2026. Every token count is either from Google's own token documentation or an illustrative assumption that is labelled as one. Model prices change often, so treat the worked examples as a method with today's inputs, not a quote, and re-run them with the AI tagging cost calculator using whatever rate your provider publishes on the day you read this.
Two-stage tagging explained
A vision model can only tag products it has been told about. The prompt therefore carries two things: the image, and a list of candidate products from your catalogue with enough detail to tell them apart. The image cost is fixed by resolution. The list cost grows with every product you include, so a 4,000-SKU catalogue pasted into every call would dwarf the image itself and also make the model worse at choosing. Two-stage tagging exists to fix that.
How a single asset gets tagged
- 01
Shortlist
Embed the post's own text (caption, detected labels) and find the nearest products by vector similarity. Small catalogues skip this and send every product.
Text only
- 02
Match
Send the image plus the shortlist to a vision model and ask for the products it can actually see, each with a confidence score, as structured JSON.
1 vision call
- 03
Validate
Throw away any product id not in the shortlist, anything below the minimum confidence, and cap the number of matches per image.
Guardrails
- 04
Route
High-confidence matches are tagged automatically. The middle band goes to a review queue. Everything else is discarded, so nothing uncertain goes live on its own.
3 bands
How Idukki runs it
Idukki's tagger follows this shape, and it is worth being exact because some older descriptions of it are out of date. It runs on Google's Gemini vision models (not Claude). For a business with up to 150 active products, every product goes into the prompt. Above that, Idukki embeds the post's caption and labels and sends the 60 closest products instead of the whole catalogue, falling back to a capped list if the post has no usable text. Where product attributes such as colour family and style have been derived, they are added to each product line so the model can separate near-identical titles. By default, matches below 0.5 confidence are dropped, matches from 0.5 up to 0.75 land in a review inbox alongside caption-based suggestions, and matches at 0.75 or above are tagged automatically. Each image can carry up to 20 matches. For video posts the tagger reads the thumbnail frame rather than every frame, which keeps video no more expensive to tag than a photo. The mechanics of the full pipeline are covered in the tagging stage of the AI UGC loop, and the product surface is on the AI tagging page.
Cost at scale: the worksheet
Per-asset cost is input tokens times the input rate, plus output tokens times the output rate. Output includes any "thinking" tokens the model spends before answering, which Google bills at the output rate, so a model that reasons at length before writing a ten-token JSON answer can cost more on the output side than the input side. Fill in the five inputs below for your own library.
| Input | How to get it | Worked-example value |
|---|---|---|
| Image tokens per asset | Google counts an image of 384 px or less on both sides as 258 tokens; larger images are tiled into 768 x 768 tiles at 258 tokens each. Confirm with the API's token-count call on a sample. | 1,032 (a 1080 x 1080 image as 4 tiles), illustrative |
| Prompt tokens per asset | Instructions plus one line per candidate product. Count a real prompt with the token-count call. | 1,200 (instructions plus a 60-product shortlist), illustrative |
| Output tokens per asset | The JSON answer is small; thinking tokens are the unknown. Log real usage on 100 assets. | 400, illustrative budget |
| Images per asset | One for photos. For video, one if you tag the cover frame, more if you sample frames. | 1 |
| Rate card | The provider's current pricing page, on the day you run the numbers. | Gemini 3.6 Flash: $0.75 in / $3.75 out per 1M tokens (standard, through 31 Dec 2026) |
With those values, one asset uses about 2,232 input tokens and 400 output tokens. On Gemini 3.6 Flash at standard rates that is roughly $0.00167 of input plus $0.0015 of output, so about $0.0032 per asset, or $3.17 per 1,000. Google's batch tier charges half the standard rate for the same model, which suits a one-off backfill where nobody is waiting on the result. Google's page also states that Gemini 3.6 Flash rises to $1.50 in and $7.50 out per million tokens from 1 January 2027, so any budget that runs into next year should use the higher figure. Gemini 3.5 Flash-Lite is cheaper again at $0.30 in and $2.50 out (standard), worth testing if its accuracy holds on your catalogue.
| Library size | 3.6 Flash, standard (2026) | 3.6 Flash, batch (2026) | 3.6 Flash, standard (from 2027) | 3.5 Flash-Lite, standard |
|---|---|---|---|---|
| 1,000 assets | $3.17 | $1.59 | $6.35 | $1.67 |
| 10,000 assets | $31.74 | $15.87 | $63.48 | $16.70 |
| 100,000 assets | $317 | $159 | $635 | $167 |
| 1,000,000 assets | $3,174 | $1,587 | $6,348 | $1,670 |
Cost per 1,000 assets, worked example
- Gemini 3.6 Flash, standard from 1 Jan 2027$6.35
- Gemini 3.6 Flash, standard (2026)$3.17
- Gemini 3.5 Flash-Lite, standard$1.67
- Gemini 3.6 Flash, batch (2026)$1.59
What actually moves the bill
- Catalogue lines in the prompt. Sending 150 products instead of 60 roughly doubles prompt tokens in this example. Shortlisting is the single biggest saving on a large catalogue.
- Image resolution. A 1080 px square is four tiles; a 384 px thumbnail is one. Products that fill the frame often tag fine from a smaller image, so test a downscaled sample before paying for full resolution.
- Thinking tokens. If your model supports a thinking budget, cap it. A matching task with a fixed list rarely needs long reasoning, and every thinking token is billed at the output rate.
- Frames per video. Tagging five sampled frames costs five images. The cover frame is usually enough for a product-led clip; multi-product try-ons may justify more.
- Re-runs. A catalogue change does not require re-tagging the whole library. Re-tag only posts whose shortlist would change, or new posts, and the steady-state cost is just your monthly intake.
The embedding step for the shortlist is small by comparison. A caption and label set is typically a few hundred text tokens, and Google lists its current embedding model at $0.20 per million text tokens, so embedding 1,000 posts of about 300 tokens each costs around six cents. It matters for architecture, not for budget.
When to keep humans in the loop
No vision model gets every product right on every photo, and nobody can give you a trustworthy universal error rate for yours: accuracy depends on how distinctive your products are, how cluttered customer photos are, and how similar your SKUs look to each other. A skincare range in near-identical bottles is harder than a furniture catalogue. The honest approach is to measure it on your own library and design the review queue around the result.
Where should this match go?
Start here
What confidence did the model return for this product?
- At or above the auto-tag threshold (0.75 by default in Idukki)
Tag automatically
The product link goes live with the post. Spot-check a small random sample each week so drift is caught early.
- If spot-checks find wrong tags: Raise the threshold for that product family, or add attribute hints that separate the look-alikes.
- Between the minimum and the auto-tag threshold
Send to review
A person accepts or rejects the suggestion. This is where most of your labour cost sits, so watch the queue size weekly.
- If reviewers accept almost everything: Lower the auto-tag threshold slightly and re-measure.
- If reviewers reject most of it: Improve the shortlist or product titles before touching thresholds.
- Below the minimum (0.5 by default in Idukki)
Discard
Not shown to anyone. If you find untagged posts that obviously show a product, tag them manually and treat them as missed recall in your next measurement.
Size the review queue before you size the model bill
The review queue is where the money goes. As an illustrative calculation: if 20% of a 100,000-asset backfill lands in review and a reviewer spends 10 seconds per suggestion, that is 20,000 decisions and about 56 hours of work. At almost any loaded hourly rate, that is several times the model bill in the table above. Halve the review share and you save more than switching to the cheapest model would. So the order of work is: shortlist well, set thresholds from evidence, and only then shop around on per-token price.
- 1Label a sample. Take 200 to 300 recent posts, have one person tag them by hand, and treat that as ground truth.
- 2Run the tagger on the same sample. Record every suggested product with its confidence score.
- 3Plot precision by band. For each 0.05 confidence band, what share of suggestions were right? The band where precision reaches the level you would accept on a live PDP is your auto-tag threshold.
- 4Count the middle. The share of assets whose best match sits between your minimum and auto-tag threshold is your review rate. Multiply by library size and seconds per decision.
- 5Repeat quarterly. New product lines and seasonal content shift the numbers, so the thresholds that were right in spring may not be right by peak season.
For the review itself, keep it fast: show the post, the suggested product image and title side by side, and make accept and reject single actions. The broader moderation workflow (brand safety, rights, sentiment) is a separate queue with different reviewers, covered in moderating UGC at scale. In Idukki, suggestions from both vision and captions surface together for accept or reject, and posts can still be tagged by hand, with hotspots, as described in the manual tagging help article.
Build it yourself or use a platform
The token bill is the smallest part of an in-house tagger. The real costs are the catalogue sync that keeps product ids current, the embedding index and its re-builds, the queueing that stops a backlog spiking your API spend in one night, the review interface, and the on-call time when a provider changes a model or a rate. None of those shows up in a per-token estimate. Idukki prices on widget impressions and includes AI tagging in every plan according to its pricing page, so there is no separate per-asset tagging bill to forecast. For the fuller build-versus-buy arithmetic, see the cost of UGC: build versus buy.
Your own pipeline
You call a vision API directly and own everything around it.
Wins at
- Full control of model, prompt and thresholds
- Token cost is transparent and yours to optimise
- Can tag assets that never go near a storefront
Struggles with
- Catalogue sync, embeddings and queueing to build and run
- A review UI to build
- Model and rate changes are your problem
- Rights and publishing still need a separate system
Tagging inside a UGC platform
Tagging runs where the posts, products and widgets already live.
Wins at
- Tags flow straight into shoppable galleries and hotspots
- Catalogue already synced for the product list
- Review queue sits next to rights and moderation
- No separate per-token bill to forecast on Idukki
Struggles with
- Less control over the model choice
- Thresholds are the platform's defaults unless it exposes them
- Tagging is only for assets in the platform
Both can work. The question is which costs you are prepared to own.
Questions CTOs ask about AI tagging cost
How much does it cost to AI-tag 10,000 UGC photos?
In the worked example here (about 2,200 input and 400 output tokens per image), roughly $32 on Gemini 3.6 Flash at standard 2026 rates, about $16 on batch, and about $63 at the rate Google publishes for that model from 1 January 2027. Your figure depends mainly on image resolution and how many catalogue products you put in each prompt.
Is batch processing worth it for tagging a back catalogue?
Usually yes. Google charges half the standard rate on its batch tier for Gemini Flash models, and a backfill has no shopper waiting on the result. Keep standard or real-time calls for new posts you want live the same day.
Does tagging video cost much more than photos?
Only if you analyse many frames. Tagging the cover frame costs the same as a photo. Sampling five frames costs five images. Idukki tags video posts from the thumbnail frame.
What confidence threshold should auto-tag products?
Set it from a labelled sample of your own posts: pick the confidence band where precision reaches what you would accept on a live product page. Idukki's defaults are 0.75 to auto-tag and 0.5 as the floor for review, which are sensible starting points rather than universal answers.
What share of assets will need human review?
It depends on your catalogue and content, so measure it: run the tagger on 200 to 300 hand-labelled posts and count how many land in your review band. That share, times library size and seconds per decision, is your labour estimate.
Which model does Idukki use for tagging?
Google Gemini vision, with Gemini embeddings to shortlist products for catalogues above 150 items. Tagging is included in Idukki plans rather than billed per token.
Sources and further reading
Send the link to your inbox.
The full worksheet is on this page. Drop your email and we’ll send you the link so you can come back to it. One email, no drip sequence.
- How two-stage tagging works: a cheap shortlist step, then one vision call per asset
- Per-asset token maths you can check against your own images and catalogue
- Worked cost projections at 1k, 10k, 100k and 1M assets, at list prices checked on 24 September 2026