Identity registry for AI training data

The identity layer the AI data supply chain is missing.

Every person in your corpus shows up a dozen ways: a name in a contract, a nickname in Slack, an email in the CRM. A7 Labs resolves them into one identity before de-identification, so the context survives all the way to the model.

Working with a small number of selected enterprise partners

The registry

Six people? Or one?Twelve identifiers. One person.

After de-identification, each name is its own PERSON token, and his email, phone and SSN are typed tokens linked to no one.One canonical identity. Names resolve to it; typed identifiers stay typed and link to it. 12 identifiers, 8,258 mentions across 149 documents.

6 people · 6 unlinked identifiers · 0 relationships 1 person · 12 identifiers · 8,258 mentions · 3 relationships
tokens_today.json registry/thomas-erickson.json

          

How training data is made today

  1. Email, Slack and Teams, contracts, CRM records, decks: terabytes of unstructured data from one enterprise, bought to become training data.

  2. Tommy in a Slack thread, Thomas Erickson in email, Mr. Erickson in an external email, his phone number in a deck, his SSN in a CRM document. De-identification turns the names into three different PERSON tokens and leaves his EMAIL, PHONE and SSN tokens linked to no one.

  3. The record that says all six are Thomas has no owner: not the seller, not the data lab, not the de-identification vendor. Every link between Thomas and his decisions, actions and outcomes is lost.

  4. Right after acquisition and before de-identification, we build the registry record: one identity, every alias and identifier confirmed, one token. PERSON_0412 is Thomas everywhere he appears.

  5. Not just Thomas. Every person, organization and deal in the corpus gets one identity and keeps its relationships. More signal per token, less hallucination.

Where we sit

Between acquiring the corpus and de-identifying it.

Everyone is racing to own dataset preparation end to end. Nobody owns the identity registry, and each player assumes someone else does.

01 · Source Enterprises & vertical AI companies Email, Slack, contracts, CRM
02 · Acquire Data labs Source and buy the corpus
A7 Labs sits hereNobody sits here
03 · A7 Labs03 · Missing Identity registry
  • Thomas EricksonPERSON_0412PERSON_017
  • TommyPERSON_0412PERSON_342
  • Mr. EricksonPERSON_0412PERSON_088
  • EMAIL_0501⇢ PERSON_0412→ nobody
  • SSN_0913⇢ PERSON_0412→ nobody

Empty.

Nothing passes through here today. The corpus goes straight from the data lab to de-identification, so nothing links Tommy to Thomas.

One identity per person, before anyone tokenizes anything.

04 · Protect De-identification One token per person A new token per mention
05 · Structure Context layer Built on intact relationships Relationships lost
06 · Train Frontier labs & enterprises AGI and open-weight models

Who builds the identity registry?

SSeller · enterprises & vertical AI companies

“The data lab will sort out who's who.”

DData lab

“The de-identification vendor handles identities.”

VDe-identification vendor

“The data lab gives that to us.”

!Result A7A7 Labs

Everyone does a piece of it. Nobody does all of it. Every mention becomes a stranger, and the links are lost.

“We build it.” The only company working 100% on the identity registry for unstructured data.

What the registry resolves

Lose the links, lose the context. Keep them, and the model learns who did what.

01

Identities

Every person and organization across the corpus, each with one consistent token.

PERSON_0412 · ORG_0091
02

Aliases

Names, nicknames, titles and handles, confirmed and mapped back to the right person.

"Tommy" → PERSON_0412
03

Identifiers & attributes

Emails, phone numbers, SSNs, employee IDs, roles and employers stay attached to one identity through de-identification.

role: "CFO" · treatment: Tokenize
04

Relationships

Who signed, approved, reported to and worked with whom, preserved across documents and systems.

signed → DEAL_0210

About

Built by the people who built enterprise AI privacy infrastructure.

We built the AI privacy product at a leading data privacy infrastructure company from the first line of code and ran it for more than three years. Our customers were the data providers that supply the frontier labs, AI-native healthcare companies, global retailers and the labs themselves.

That is where we saw the gap. Every one of them acknowledges it. None of them owns it.

  • Volume. Tens of terabytes of unstructured data a month through production NER models, for leading AI companies, every month.
  • Depth. Years of day-and-night work on entity detection, tokenization and consistency across real enterprise corpora.
  • Proximity. We sat inside the pipelines that prepare data for the frontier labs and watched this exact step break.

The only company working 100% on the identity registry for unstructured data. We waited for the right problem. This is it.

Work with us

Send us a corpus. We'll send back its people.

We are currently working with a small number of selected enterprise partners and data labs. Tell us about your data and we'll run a sample through the registry.

  • Data labs and dataset builders
  • De-identification and privacy vendors
  • Enterprises training open-weight models
I'm interested in *