AI Entity Enrichment PlatformYour documents already describe a data model.
Entity Enricher reads a PDF, a photo or a voice memo, designs the schema, fills it from those documents, the web and several LLMs, and returns validated, structured records — kept as JSON, or synced to a database you own.
Run a live enrichmentAn enrichment starts from a request and fills the schema for one entity and those linked to it. Five players are enriched in turn: a club or a country found for an earlier player is recognized, not added again.
Every Field of a Schema Is a Question
The enrichments above fill a schema with values. Here is the same schema as a class diagram: each object of its JSON structure, with its fields, each showing its JSON type (string, number, boolean…) where a value would stand. Every field carries a description, and resting the pointer on an object, or tapping it, opens the descriptions of its fields. A link to another object is a field too: it carries the field's name and opens the question that tells the model which object it links to.
Each description is the question the language model is asked for that field's value. Schema generation writes them, and the schema editor can edit them, to make a question more precise or to remove an ambiguity before any enrichment runs.
One Identifier per Club with semantic IDs, However Its Name Is Written
Each enrichment writes a club the way its sources do: FC Barcelona in one player's record, Futbol Club Barcelona in the next. Every club a record names gets a semantic ID, never filled by the model, resolved in the three steps below, cheapest first. When in doubt, a new ID is created: merging two IDs takes one click, undoing a wrong merge does not.
A Schema Built by AI From Your Data Samples
A schema can be generated from a sample of your data or from a document, written field by field in the schema editor, or both: a generated schema is a first draft, edited afterwards like any other. The steps below follow its generation from a sample.
From a sample to an entity map
Entity Enricher reads a sample of your data and recognizes the entities it describes — here, books, authors, series and publishers — along with the fields that identify each one. These entities are proposals: the schema keeps the sample's flat shape until the user turns one into an entity of its own, in one click. Its fields then move under it, and an author cited by several books becomes one entity with its own unique identifier.
Fields shared out among experts
Schema generation groups the fields by the knowledge they call for, and gives each group an expert: here a literary expert for titles, genres, authors and series, and a publishing expert for publishers, prices and publication years. Fields are shared out one by one, not entity by entity: a book's title goes to one expert, its price to the other. A schema gets at most one expert for every six fields, so a small one is not split too finely.
At enrichment, each expert can become a call of its own, run in parallel: a model asked only about prices and publishers answers them more fully than one asked about everything at once.
Each expert describes its fields
Each expert then writes the description of its own fields, in its own vocabulary: which year a book's year is, in which currency and at which date its price is read. The experts write side by side, one call each. A description is the question a model answers for its field at enrichment, so a precise one is what makes two models, or two runs, fill the field alike.
Fields that identify an entity
For each entity type, schema generation picks the fields that identify one instance — a book's title, an author's name. Dante and Dan Brown each wrote an Inferno: a book is told apart by its title and its author together.
A semantic ID for each entity
When semantic IDs are requested, each object that represents an entity gets an id field: the book, and each entity nested under it — author, series, publisher. The model never fills it: after each enrichment, it is resolved from the entity's identifying fields, so the same author written two ways keeps one identifier.
An object that only holds facts of its parent — here the book's year and price — gets none: no field in it identifies anything. Neither does a pairing, such as a book's volume number in its series: it is identified by the two entities it relates.
Multilingual fields
A text field can be marked multilingual: it then holds one value per requested language — a title, a genre or a country written in each of them, a name in each language's script. A year or a price stays a single value. Identity is read in the first language, so a translated title never creates a second book.
Identifiers kept as they are
For each entity, one or several language models fill the schema from the attached documents and from their own knowledge. A field marked preserved is the exception: its value comes from your own system — a row id, an internal reference — and passes through the enrichment unchanged, so each enriched book goes back to the row it came from. Unmarked, the field is the model's to fill, and nothing keeps it from writing the ISBN it knows in place of your id.
From Raw Data to Information System
One pipeline takes whatever you have — documents, spreadsheets, half-filled rows — and returns records your database can trust.
Source
Bring a batch from your existing system — or a single new entity the moment it appears. Documents, images, web search, and LLM world knowledge fill what your data doesn’t say.
Structure
Describe your target in plain language or paste a sample — AI drafts a typed schema with expertise domains. Refine it visually or by chat.
Verify
Multiple models answer in parallel, per knowledge domain. Conflicts are detected field by field and resolved by rules or an AI arbiter — with the reasoning recorded.
Integrate
Validated records flow back with your original keys preserved verbatim and semantic IDs as stable join keys. No duplicates, no re-keying — up to 40 languages per field.
Two ways to feed it
Batch — from your existing system
Pull hundreds of entities from your database, CRM, or any REST endpoint — paste JSON or fetch a URL with auth. Enrich them in parallel, watch progress live, write clean records back — or export to Excel.
On the fly — as new entities arrive
A new lead, product, or document enters your system? Enrich it in seconds — one API call, an n8n/Make trigger, or straight from a chat via MCP. Structured, validated, ready to insert.
Both paths share the same semantic IDs — an entity enriched today in a batch and re-encountered tomorrow on the fly still lands on one record.
Every Value Has a Paper Trail
Most AI tools ask you to trust the output. We let you inspect how it was decided.
Pre-flight check
A fast model classifies the entity against your schema first. Enriching “Titan” as a planet? You’re warned before a single token is spent.
Models in competition
Two or more LLMs answer independently. Outputs are schema-validated; errors go back to the model to self-correct, automatically.
Arbitrated, recorded
Field-level conflicts resolved by majority, median, or an AI arbiter. Every decision — all candidate values, the winner, the reasoning — is stored on the record.
Acme Corp
Any entity: company, drug, legal case, research paper...
Match — Company
Catches type mismatches before wasting LLM credits.
Bring your own API keys — works with any LLM provider.
Schema split by domain — self-correcting prompts retry on validation failure.
Deep merge of expertise responses per model.
Acme Corp
ArbitratedReasoned field-level conflict resolution produces the final trusted result.
Eight defense layers stand between an LLM’s imagination and your database. How we prevent hallucinations →
Your Data, Your Models, Your Keys
Built for teams whose data can’t leave the building — cloud convenience with full control over where inference happens.
Bring your own API keys
Use your organization’s Anthropic, OpenAI, or Gemini keys — your billing, your data-processing agreements. Platform keys are just the zero-setup default.
Run models on your own hardware
Pair a laptop or on-prem GPU server in two minutes and route enrichments through a secure tunnel to your local Ollama. Sensitive data never reaches a cloud LLM.
How the tunnel works →Tenant-isolated by construction
Records, schemas, files, and concept registries are organization-scoped. Role-based access control down to each API key.
Organizations & roles →Why Entity Enricher
Purpose-built for structured LLM enrichment — here is what it adds over a single model call or a hand-built pipeline.
Plugs Into Your Information System
Built-inAdd a database sync to a schema and enrichments mirror into your own PostgreSQL as real relational tables — SQL snapshot to seed it, an idempotent delta feed to keep it converged, and an open-source sync client that applies it with zero glue code. Input keys are preserved verbatim, and each entity gets a stable semantic ID: a ready-made join key that resolves “Headache”, “Céphalée” and “Cephalalgia” into one record, not three.
Raw JSON per call. You design storage, keys, matching, and reconciliation before it can touch your systems.
Custom Schema
You define the output structure. Any entity type, any fields, any nesting depth.
One hand-rolled prompt returns one flat shape. Nesting, types, and validation are all yours to build.
Multi-Model
Run 2+ LLMs simultaneously. Compare results. Use the best of each.
Single model, single provider. No way to cross-validate or improve accuracy.
Fusion & Arbitration
Field-level conflict detection with rule-based or LLM-arbitrated resolution.
Blind trust in a single source. No conflict awareness.
Any Domain, Any Entity
Legal entities, pharma compounds, research papers, real estate — anything.
Every new domain means new prompts, new validation, new plumbing — a pipeline per entity type.
Multilingual by Design
Built-inMark a field multilingual once. A single enrichment call returns the value translated into every language you selected — up to 40 — with no extra LLM calls or translation pipeline.
English-only output. Translation is a separate step, a separate cost, and a separate failure mode.
Bring Your Own Documents
NewAttach PDFs, slides, spreadsheets, contracts, scans, audio recordings. Vision-, PDF- and audio-capable models read them directly; the rest are extracted server-side and inlined automatically.
Text-only inputs. Documents are your problem — convert, OCR, transcribe, chunk, and clean before you can enrich.
Cost-Optimized by Default
Built-inPrompt caching reuses the shared prompt across parallel calls at ~10% of input price, each expertise only sees its own fields, and a cheap pre-flight check stops you paying to enrich the wrong entity.
Flat per-record pricing with no token-level optimization — and no visibility into what you actually spent.
Works Where You Work
Design your schema once, then enrich at any scale — from the web app, from automated workflows, or straight from your own code.
Batch Enrichment
Enrich hundreds of entities in parallel from the web app. Real-time streaming, auto-fusion, Excel export.
n8n & Make Workflows
Automated pipelines: trigger on new data, enrich, push to your CRM or database. 400+ app integrations.
REST API
Programmatic access for custom integrations. Typed OpenAPI schema, org-scoped keys, sync and streaming endpoints.
Connect to 400+ apps via n8n
Build automated enrichment pipelines with n8n's visual workflow editor. Pull data from any source, enrich with AI, and push results anywhere.
Entity Enricher ships an embedded MCP (Model Context Protocol) server. List your schemas, enrich an entity, inspect the result — all from the chat. No workflow editor required.
The open-source ee-database client keeps your own PostgreSQL converged with every enrichment — snapshot bootstrap, then a live delta feed over an outbound connection. No workflow to build, and your connection string never leaves your machine.
Built for Every Domain
Not just B2B contacts. Define a schema for any entity type and enrich it.
How We Compare
Coming from a web-research API, a spreadsheet research agent — or building your own LLM pipeline? Here’s where Entity Enricher stands.
| Feature | Entity Enricher | Parallel AI | AI Research Agents | DIY LLM Pipeline |
|---|---|---|---|---|
| Custom Nested Schema | Flat JSON schema | One prompt per column | Hand-coded | |
| Multi-Model Enrichment | You orchestrate | |||
| Fusion & Conflict Resolution | ||||
| Field-Level Audit Trail | Citations + confidence | |||
| Linked-Entity Dedup (Semantic IDs) | ||||
| Sync to Your Own PostgreSQL | CSV / sheet export | You build it | ||
| Your Documents as Sources | Varies | You build it | ||
| Any Entity Type | ||||
| Multilingual Output (40 Languages) | You build it | |||
| BYOK / Self-Hosted Models | Rarely | |||
| Batch Processing | You build it | |||
| Maintenance | Managed | Managed | Managed | Yours, forever |
| Pricing | Pay-per-token | $5-25 / 1k rows | Credits / subscription | Eng time + tokens |
Backfill the past. Enrich the future.
Your company’s knowledge is already written down — make it queryable. Start free, bring your own API keys, and pay only LLM costs.
Get Started Free