The Semantic IDs page

Every enrichment that resolves a semantic ID either reuses a concept your organization already knows or mints a new one. The Semantic IDs page is where that growing vocabulary becomes something you can look at: browse it, measure how close two entries really are, add terms by hand, answer the questions the identity judge leaves for you, retire what you no longer want, and move the whole thing to a different embedding model.

Opening the page

It lives at /semantic-ids in the sidebar, and it is only useful once your organization has an embedding model and at least one schema carrying a semantic ID — concepts are created by enrichments, not by the page.

Who can do what

  • •Editor — browse, compare, export, check a text, add a concept, delete concepts.
  • •Owner / admin — also create concepts through an import (opening a new concept type with them, if that is what the file needs), clear whole concept types, and migrate the embedding model.

Reading the concept table

Each row is one concept: its canonical text (the identity text that first created it), its concept type (the space it lives in — the entity type name by default), how many records use it, and when it was created. Filter by text, by any number of concept types, by the schemas that write into them, or by a minimum usage to find the entries worth your attention. Picking schemas is a shortcut for their concept types — and since a type is shared, it also shows the concepts another schema created in that same type. Export view downloads exactly what the filters are showing.

  1. 1Alias count on the canonical row
  2. 2A spelling absorbed by the identity judge
The +1 chip is the alias count: one concept, several spellings, one usage total. A form the judge absorbed carries auto — the judge is paid once per surface form, ever.
Selecting a row turns the Similarity column into a measurement against it, and opens that concept on the right.

Similarity is measured inside one space only. Vectors are comparable only within the same concept type and embedding model. Rows outside the selected concept’s space show —, which means “not comparable” — never “0%”.

Adding terms: the curate loop

Sometimes you know the vocabulary before the data arrives — a list of dishes, statuses, product families. The Curate bar is built for typing that list in: scope it to a concept type once — the scope goes into the page address, so /semantic-ids/<concept type> is a link straight to that slice — then repeat type a value → Check → Add. Enter checks, a second Enter adds, and the field clears and keeps focus, so a fifty-term vocabulary is a few minutes of typing without touching the mouse.

Nothing resolves this text above the threshold, so Add concept is armed. The table now ranks every concept by its similarity to what you typed.

Check is a dry run of the real resolution ladder — the same one an enrichment runs — so its verdict is not an estimate. And because you have already paid for that verdict, it doubles as the duplicate guard: when an existing concept covers your text at or above the threshold, Add concept stays disabled.

An exact hit costs nothing at all — no model call. A near match reports the incumbent and its score instead.

That refusal is deliberate: a twin above the threshold could never win a resolution, and would split future matches unpredictably between the two entries. The Threshold slider next to the field decides where that line sits for your checks — raise it to be stricter, lower it to merge more aggressively.

The vocabulary does not have to start from an enrichment either: New concept type… (in the page toolbar, editor and up) names a type and picks the embedding model its concepts will live in — by default your organization’s model. The curate bar is scoped to it right away, and the type joins the vocabulary with the first concept you add; a type that already exists keeps its model, since moving a vocabulary between models is what the migration is for.

The embedding model is the one question that cannot be asked later: a type keeps the model it was opened with, and moving it is a migration.

The review queue

Nothing lands in Review for merely looking similar. Every pair here is something the identity judge decided and left for you: it could not tell (Unsure), it found two of your concepts equal to one incoming text (Looks duplicated), or it separated two that measure almost identical (Kept apart) — that last one is offered so you can confirm it did not overlook the obvious candidate. Each row carries the judge’s own sentence explaining what decided it.

Two very different verdicts side by side: two names for one turbot dish, and two genuinely different chocolate desserts. The judge finds the questions; you answer them.

Compare → jumps back to the table with that concept selected, so you can see what else sits near it before deciding. Merge… folds one into the other — the loser’s spellings become aliases of the winner, so the pair converges everywhere and stays converged. Dismiss closes the question when the two really are two things. Usage counts tell you which of the two the data actually prefers.

The impact is counted before you confirm: records re-pointed, and the database changes the merge queues downstream.

Inspecting one concept

The right-hand panel shows the concept’s ID (copyable — it is what your database joins on), its normalized text, the embedding model behind it, and the records that resolved to it. The selected concept is part of the page address, so the URL in your address bar is a link straight back to it — shareable with a colleague, or worth keeping in a ticket. Two views place it among its neighbours.

Orbit — distance from the centre is similarity, so the dashed ring is the threshold itself: anything inside it would resolve to this concept.
Concept map (3D) — an opt-in layout of the whole space. Colour is the measurement; position only hints at cluster shape, which is why no threshold sphere is drawn.

Why position isn’t truth

Flattening 1536 dimensions into 3 cannot preserve distances, and the layout exaggerates how tight clusters are. Both views paint every point from the same similarity scale, so read the colour, not the gap. The first map of a session takes a few seconds to lay out; later ones are instant.

Linked records

Each linked record shows the score at which it resolved here. A — means the ID arrived in the input and was passed through, so nothing was ever compared — the free, unambiguous path described in the semantic-ID guide.

Importing and exporting a vocabulary

Import & resolve runs a whole CSV column through the same ladder and annotates every row with what would happen to it — there is no file-size limit; large files resolve in batches of 1000 with live progress. The file is parsed in your browser — only the values themselves are sent.

A match-only run writes nothing, so it is a faithful preview of what your next enrichment would do with the same values.
OutcomeWhat it means
exactThe same text, normalized, already exists — free, no model call.
matchedA different wording resolved to an existing concept above the threshold.
would_mintNothing was close enough; an enrichment would create a new concept here.
mintedThe same case, with creation switched on — the concept now exists (owner only).

The concept type is typed, not picked. A file is a perfectly normal way to start a vocabulary, so the field suggests the spaces you already have while accepting a name you do not — and saysnew when that is what it is. A new name asks the one question that cannot be asked later: which embedding model its concepts will live in. Answer it here and the space is created with the first value the import mints; afterwards, moving a vocabulary between models is a migration.

Two things keep a slip from becoming a second vocabulary: typing the name of a space you already have selects it, whatever the capitalisation, and a name one or two letters away from an existing one is called out with a one-click correction. Against a genuinely new space there is nothing to compare with, so a match-only run answers “every value would be created” without embedding a single row — and charges nothing for it.

The resolved rows download as their own CSV, so an import can also be used purely as an audit: which of our 900 supplier names are already known, and which would open a new identity? Exports work the other way and name the file after the filters they were taken with, so a folder of them stays self-describing.

Deleting concepts

Deleting is safe in a way most data deletions are not: the vocabulary rebuilds itself, because the next enrichment simply mints what it needs again. What does not come back is convergence with the IDs you already stored, and the dialog says so with real counts before you confirm.

Each type names the embedding space it lives in — two types on different models never compare, which is why the count is broken down this way.
Records keep the ID they already hold; new enrichments of the same thing will mint a different one. Downstream databases see that as a new row.

After such a change, keeping the old concepts is the risky choice, not the cautious one: identities composed from the new keys can still land within the threshold of the old vectors and be quietly absorbed by them, leaving you with IDs that mean neither thing.

Changing the embedding model

Vectors from two different models are not comparable, so switching model means re-embedding every concept. Once concepts exist, this migration is the only sanctioned way to do it — and it is the reason the organization’s embedding-model setting refuses to change on its own.

One embedding space at a time. Concept IDs never change, so records, entities, exports and database syncs stay valid throughout.

Dry run first

The preview reports what the re-embedding will cost, and which concept pairs are likely to collide — land within the threshold of each other in the new space and start resolving together. It measures the realistic candidates rather than every pair, and says so.

No pause, no window to babysit

Enrichments keep resolving in the old space while the new vectors are built beside it; concepts minted meanwhile are picked up by a later pass. The switch happens in a single step at the end, which also flips your organization’s default embedding model. Interrupting it is harmless — starting again continues where it stopped.

What the page costs

Checks, adds and imports are billed like any other embedding usage, and exact hits cost nothing because they never reach a model. Browsing, comparing, the orbit, the 3D map, exporting and deleting are free. A day’s worth of interactive spend is folded into a single line on your credit history, so the ledger stays readable instead of filling up with fractions of a cent.