AI Entity Enrichment PlatformYour documents already describe a data model.

Entity Enricher reads a PDF, a photo or a voice memo, designs the schema, fills it from those documents, the web and several LLMs, and returns validated, structured records — kept as JSON, or synced to a database you own.

Run a live enrichment

An enrichment starts from a request and fills the schema for one entity and those linked to it. Five players are enriched in turn: a club or a country found for an earlier player is recognized, not added again.

Every Field of a Schema Is a Question

The enrichments above fill a schema with values. Here is the same schema as a class diagram: each object of its JSON structure, with its fields, each showing its JSON type (string, number, boolean…) where a value would stand. Every field carries a description, and resting the pointer on an object, or tapping it, opens the descriptions of its fields. A link to another object is a field too: it carries the field's name and opens the question that tells the model which object it links to.

Each description is the question the language model is asked for that field's value. Schema generation writes them, and the schema editor can edit them, to make a question more precise or to remove an ambiguity before any enrichment runs.

One Identifier per Club with semantic IDs, However Its Name Is Written

Each enrichment writes a club the way its sources do: FC Barcelona in one player's record, Futbol Club Barcelona in the next. Every club a record names gets a semantic ID, never filled by the model, resolved in the three steps below, cheapest first. When in doubt, a new ID is created: merging two IDs takes one click, undoing a wrong merge does not.

A Schema Built by AI From Your Data Samples

A schema can be generated from a sample of your data or from a document, written field by field in the schema editor, or both: a generated schema is a first draft, edited afterwards like any other. The steps below follow its generation from a sample.

  1. From a sample to an entity map

    Entity Enricher reads a sample of your data and recognizes the entities it describes — here, books, authors, series and publishers — along with the fields that identify each one. These entities are proposals: the schema keeps the sample's flat shape until the user turns one into an entity of its own, in one click. Its fields then move under it, and an author cited by several books becomes one entity with its own unique identifier.

  2. Fields shared out among experts

    Schema generation groups the fields by the knowledge they call for, and gives each group an expert: here a literary expert for titles, genres, authors and series, and a publishing expert for publishers, prices and publication years. Fields are shared out one by one, not entity by entity: a book's title goes to one expert, its price to the other. A schema gets at most one expert for every six fields, so a small one is not split too finely.

    At enrichment, each expert can become a call of its own, run in parallel: a model asked only about prices and publishers answers them more fully than one asked about everything at once.

  3. Each expert describes its fields

    Each expert then writes the description of its own fields, in its own vocabulary: which year a book's year is, in which currency and at which date its price is read. The experts write side by side, one call each. A description is the question a model answers for its field at enrichment, so a precise one is what makes two models, or two runs, fill the field alike.

  4. Fields that identify an entity

    For each entity type, schema generation picks the fields that identify one instance — a book's title, an author's name. Dante and Dan Brown each wrote an Inferno: a book is told apart by its title and its author together.

  5. A semantic ID for each entity

    When semantic IDs are requested, each object that represents an entity gets an id field: the book, and each entity nested under it — author, series, publisher. The model never fills it: after each enrichment, it is resolved from the entity's identifying fields, so the same author written two ways keeps one identifier.

    An object that only holds facts of its parent — here the book's year and price — gets none: no field in it identifies anything. Neither does a pairing, such as a book's volume number in its series: it is identified by the two entities it relates.

  6. Multilingual fields

    A text field can be marked multilingual: it then holds one value per requested language — a title, a genre or a country written in each of them, a name in each language's script. A year or a price stays a single value. Identity is read in the first language, so a translated title never creates a second book.

  7. Identifiers kept as they are

    For each entity, one or several language models fill the schema from the attached documents and from their own knowledge. A field marked preserved is the exception: its value comes from your own system — a row id, an internal reference — and passes through the enrichment unchanged, so each enriched book goes back to the row it came from. Unmarked, the field is the model's to fill, and nothing keeps it from writing the ISBN it knows in place of your id.

From Raw Data to Information System

One pipeline takes whatever you have — documents, spreadsheets, half-filled rows — and returns records your database can trust.

1

Source

Bring a batch from your existing system — or a single new entity the moment it appears. Documents, images, web search, and LLM world knowledge fill what your data doesn’t say.

2

Structure

Describe your target in plain language or paste a sample — AI drafts a typed schema with expertise domains. Refine it visually or by chat.

3

Verify

Multiple models answer in parallel, per knowledge domain. Conflicts are detected field by field and resolved by rules or an AI arbiter — with the reasoning recorded.

4

Integrate

Validated records flow back with your original keys preserved verbatim and semantic IDs as stable join keys. No duplicates, no re-keying — up to 40 languages per field.

Two ways to feed it

Batch — from your existing system

Pull hundreds of entities from your database, CRM, or any REST endpoint — paste JSON or fetch a URL with auth. Enrich them in parallel, watch progress live, write clean records back — or export to Excel.

On the fly — as new entities arrive

A new lead, product, or document enters your system? Enrich it in seconds — one API call, an n8n/Make trigger, or straight from a chat via MCP. Structured, validated, ready to insert.

Both paths share the same semantic IDs — an entity enriched today in a batch and re-encountered tomorrow on the fly still lands on one record.

Every Value Has a Paper Trail

Most AI tools ask you to trust the output. We let you inspect how it was decided.

Before

Pre-flight check

A fast model classifies the entity against your schema first. Enriching “Titan” as a planet? You’re warned before a single token is spent.

During

Models in competition

Two or more LLMs answer independently. Outputs are schema-validated; errors go back to the model to self-correct, automatically.

After

Arbitrated, recorded

Field-level conflicts resolved by majority, median, or an AI arbiter. Every decision — all candidate values, the winner, the reasoning — is stored on the record.

Entity

Acme Corp

Any entity: company, drug, legal case, research paper...

Pre-flight Classification

Match — Company

Catches type mismatches before wasting LLM credits.

Anthropic
OpenAI
Google Gemini

Bring your own API keys — works with any LLM provider.

Anthropic
FinancialsLLM prompt
LegalLLM prompt
MarketLLM prompt
OpenAI
FinancialsLLM prompt
LegalLLM prompt
MarketLLM prompt
Gemini
FinancialsLLM prompt
LegalLLM prompt
MarketLLM prompt

Schema split by domain — self-correcting prompts retry on validation failure.

Anthropic Result
OpenAI Result
Gemini Result

Deep merge of expertise responses per model.

Final Enriched Result

Acme Corp

Arbitrated

Reasoned field-level conflict resolution produces the final trusted result.

Eight defense layers stand between an LLM’s imagination and your database. How we prevent hallucinations →

Your Data, Your Models, Your Keys

Built for teams whose data can’t leave the building — cloud convenience with full control over where inference happens.

Bring your own API keys

Use your organization’s Anthropic, OpenAI, or Gemini keys — your billing, your data-processing agreements. Platform keys are just the zero-setup default.

Run models on your own hardware

Pair a laptop or on-prem GPU server in two minutes and route enrichments through a secure tunnel to your local Ollama. Sensitive data never reaches a cloud LLM.

How the tunnel works →

Tenant-isolated by construction

Records, schemas, files, and concept registries are organization-scoped. Role-based access control down to each API key.

Organizations & roles →

Why Entity Enricher

Purpose-built for structured LLM enrichment — here is what it adds over a single model call or a hand-built pipeline.

Plugs Into Your Information System

Built-in

Add a database sync to a schema and enrichments mirror into your own PostgreSQL as real relational tables — SQL snapshot to seed it, an idempotent delta feed to keep it converged, and an open-source sync client that applies it with zero glue code. Input keys are preserved verbatim, and each entity gets a stable semantic ID: a ready-made join key that resolves “Headache”, “Céphalée” and “Cephalalgia” into one record, not three.

Raw JSON per call. You design storage, keys, matching, and reconciliation before it can touch your systems.

Custom Schema

You define the output structure. Any entity type, any fields, any nesting depth.

One hand-rolled prompt returns one flat shape. Nesting, types, and validation are all yours to build.

Multi-Model

Run 2+ LLMs simultaneously. Compare results. Use the best of each.

Single model, single provider. No way to cross-validate or improve accuracy.

Fusion & Arbitration

Field-level conflict detection with rule-based or LLM-arbitrated resolution.

Blind trust in a single source. No conflict awareness.

Any Domain, Any Entity

Legal entities, pharma compounds, research papers, real estate — anything.

Every new domain means new prompts, new validation, new plumbing — a pipeline per entity type.

Multilingual by Design

Built-in

Mark a field multilingual once. A single enrichment call returns the value translated into every language you selected — up to 40 — with no extra LLM calls or translation pipeline.

English-only output. Translation is a separate step, a separate cost, and a separate failure mode.

Bring Your Own Documents

New

Attach PDFs, slides, spreadsheets, contracts, scans, audio recordings. Vision-, PDF- and audio-capable models read them directly; the rest are extracted server-side and inlined automatically.

Text-only inputs. Documents are your problem — convert, OCR, transcribe, chunk, and clean before you can enrich.

PDFPNGJPEGMP3WAVM4ADOCXDOCODTRTFEPUBHTMLCSVXLSXPPTXTXTMDSee all formats →

Cost-Optimized by Default

Built-in

Prompt caching reuses the shared prompt across parallel calls at ~10% of input price, each expertise only sees its own fields, and a cheap pre-flight check stops you paying to enrich the wrong entity.

Flat per-record pricing with no token-level optimization — and no visibility into what you actually spent.

How cost optimization works →

Works Where You Work

Design your schema once, then enrich at any scale — from the web app, from automated workflows, or straight from your own code.

Batch Enrichment

Enrich hundreds of entities in parallel from the web app. Real-time streaming, auto-fusion, Excel export.

n8n & Make Workflows

Automated pipelines: trigger on new data, enrich, push to your CRM or database. 400+ app integrations.

REST API

Programmatic access for custom integrations. Typed OpenAPI schema, org-scoped keys, sync and streaming endpoints.

Connect to 400+ apps via n8n

Build automated enrichment pipelines with n8n's visual workflow editor. Pull data from any source, enrich with AI, and push results anywhere.

Google Sheets
Source Data
Entity Enricher
Entity Enricher
AI Enrichment
HubSpot
CRM Sync
HubSpot
CRM
Salesforce
CRM
Google Sheets
Spreadsheet
Airtable
Database
Slack
Messaging
PostgreSQL
Database
Webhook
API
Gmail
Email
Notion
Workspace
Stripe
Payments
Jira
Project Mgmt
HTTP Request
API
CRM Sync
Push enriched data directly to HubSpot, Salesforce, or any CRM
Waterfall Enrichment
Chain multiple enrichment steps with conditional logic
No-Code Workflows
Visual drag-and-drop pipeline builder — no coding required
Automated Pipelines
Trigger enrichment on new rows, form submissions, or schedules
Or use it directly from Claude Desktop, Claude Code, or Cursor

Entity Enricher ships an embedded MCP (Model Context Protocol) server. List your schemas, enrich an entity, inspect the result — all from the chat. No workflow editor required.

Or mirror everything into your own database

The open-source ee-database client keeps your own PostgreSQL converged with every enrichment — snapshot bootstrap, then a live delta feed over an outbound connection. No workflow to build, and your connection string never leaves your machine.

How We Compare

Coming from a web-research API, a spreadsheet research agent — or building your own LLM pipeline? Here’s where Entity Enricher stands.

FeatureEntity EnricherParallel AIAI Research AgentsDIY LLM Pipeline
Custom Nested SchemaFlat JSON schemaOne prompt per columnHand-coded
Multi-Model EnrichmentYou orchestrate
Fusion & Conflict Resolution
Field-Level Audit TrailCitations + confidence
Linked-Entity Dedup (Semantic IDs)
Sync to Your Own PostgreSQLCSV / sheet exportYou build it
Your Documents as SourcesVariesYou build it
Any Entity Type
Multilingual Output (40 Languages)You build it
BYOK / Self-Hosted ModelsRarely
Batch ProcessingYou build it
MaintenanceManagedManagedManagedYours, forever
PricingPay-per-token$5-25 / 1k rowsCredits / subscriptionEng time + tokens

Backfill the past. Enrich the future.

Your company’s knowledge is already written down — make it queryable. Start free, bring your own API keys, and pay only LLM costs.

Get Started Free