article · featured

OpenCorporates vs Aleph: Which One Fits Which Research Job?

Choose OpenCorporates when legal-entity identity and provenance are the main uncertainty; choose OCCRP Aleph when a known entity needs cross-dataset, document and relationship context.

published
Apr 21, 2026
updated
Aug 20, 2026
slug
opencorporates-vs-aleph-which-one-fits-which-research-job
status
Published

OpenCorporates vs OCCRP Aleph: Which One Fits Which Research Job?

OpenCorporates and OCCRP Aleph are both useful for company research.

They are also easy to misuse when they are treated as interchangeable.

Both can contain companies.

Both can help resolve identities.

Both can expose relationships.

Both can be used in larger investigative workflows.

But their center of gravity is different.

A useful starting distinction is:

OpenCorporates is strongest when you need a structured legal-entity anchor with explicit provenance.

Aleph is strongest when you need to search, organize and cross-reference entities and documents across multiple datasets inside an investigative workspace.

That distinction is not absolute.

OpenCorporates contains more than bare registry rows.

Aleph can contain company registries and can support entity reconciliation.

The point is not to force each platform into one box.

The point is to understand which research problem each one handles most naturally.


Start with the research question

Do not begin with:

Which tool is better?

Begin with:

What uncertainty am I trying to reduce?

Typical company-research questions include:

Identity questions

  • Does this legal entity exist?
  • What is its exact registered name?
  • Which jurisdiction is it registered in?
  • What is the company number?
  • Is it active or dissolved?
  • Which official source underlies the record?

These are usually entity-anchor questions.

Context questions

  • Which documents mention this entity?
  • Does it appear across multiple datasets?
  • Which people, companies, contracts or assets are connected to it?
  • Is it mentioned in leaks, procurement data, court records or sanctions lists?
  • Can I compare a list of entities against a broader corpus?

These are usually context-expansion questions.

Investigation-management questions

  • Can I upload my own documents?
  • Can I build a case workspace?
  • Can I draw relationships?
  • Can I organize timelines?
  • Can I cross-reference my own entities against available datasets?

These are investigation-workspace questions.

The right tool depends on which class of question you are asking now.


OpenCorporates: structured company identity first

OpenCorporates describes itself as a large open database of companies built from primary public sources.

Its API documentation emphasizes:

  • structured company data;
  • source and provenance information;
  • company identifiers;
  • jurisdiction;
  • officers;
  • filings and related public-record information where available;
  • reconciliation of company names to legal corporate entities.

This makes it particularly strong when your first job is:

stabilize the legal entity before expanding the investigation.


Why provenance matters in OpenCorporates

A useful company record should not be a detached assertion.

You want to know:

  • where the information came from;
  • which publisher supplied it;
  • when it was retrieved;
  • whether it came from an external public source or another process.

OpenCorporates' API documentation exposes provenance/source information specifically for this reason.

That is analytically important.

Compare these two findings.

Weak:

OpenCorporates says Example Ltd is active.

Stronger:

OpenCorporates records Example Ltd, company number X in jurisdiction Y, and provides provenance to source Z, retrieved at time T.

The second statement preserves the path back to the underlying evidence.

For important findings, follow that provenance to the primary registry when practical.


OpenCorporates is especially useful for entity normalization

Company names are messy.

The same entity can appear as:

Example Holdings Limited
EXAMPLE HOLDINGS LTD
Example Holdings, Ltd.
Example Holdings

and different entities can share nearly identical names.

OpenCorporates provides reconciliation capabilities designed to help match company names to legal corporate entities.

Its OpenRefine reconciliation API is specifically intended for that use case.

This is useful when you start with:

  • a spreadsheet of companies;
  • names extracted from documents;
  • supplier lists;
  • counterparties;
  • messy legacy data.

The objective is not:

find a similar-looking company.

It is:

identify the most plausible legal entity and retain the identifiers that let you distinguish it from others.


What OpenCorporates does not automatically solve

A structured legal-entity record does not answer every investigative question.

OpenCorporates does not automatically tell you:

  • why an entity appears in a leaked document;
  • whether a contract relationship is suspicious;
  • what a court filing means;
  • which documents in a large corpus mention the entity;
  • how a person, company, asset and contract relate inside your specific investigation;
  • whether a name match in another dataset is genuinely the same entity.

Even strong registry data still needs context.

That is where a platform like Aleph becomes useful.


OCCRP Aleph: multi-dataset and documentary context

Aleph is designed around datasets, documents and entities.

Its underlying FollowTheMoney model can represent many entity types, including:

  • people;
  • companies;
  • assets;
  • contracts;
  • events;
  • bank accounts;
  • addresses;
  • vessels and other structured objects.

Aleph datasets can include categories such as:

  • company registries;
  • sanctions lists;
  • procurement;
  • court archives;
  • regulatory filings;
  • leaks;
  • document libraries;
  • financial records;
  • news archives.

That means Aleph is not merely a document search engine.

It is a system for bringing heterogeneous evidence into a common investigative model.


Cross-referencing is one of Aleph's defining strengths

Aleph's documentation describes cross-referencing as comparing entities from one dataset against entities in other datasets to find leads and patterns.

The critical detail is that good cross-referencing depends on keys that help distinguish one entity from another.

For a company, useful keys can include:

  • registration number;
  • jurisdiction;
  • address;
  • other identifiers;
  • structured attributes.

This is the same principle used in disciplined entity resolution:

names generate candidates; identifiers reduce ambiguity.

Aleph therefore can contribute to entity resolution too.

The difference is that its main strength is often what happens after you have a useful entity representation:

compare it across other datasets and documents.


Aleph becomes more valuable as the investigation gets richer

Imagine you have already established:

Legal name: Example Holdings Ltd
Jurisdiction: GB
Company number: 12345678

Now you want to know:

  • Does this entity appear in procurement datasets?
  • Does it appear in a leak?
  • Are documents connected to known officers?
  • Do related entities appear elsewhere?
  • Can I upload my own files and cross-reference them?
  • Can I build a relationship graph or timeline?

This is where Aleph's model fits naturally.

The legal identity is the anchor.

Aleph helps turn the anchor into investigative context.


Aleph investigation workspaces are a different class of capability

Aleph supports investigation workspaces where users can upload and organize:

  • entities;
  • documents;
  • email archives;
  • spreadsheets and other files.

Its documentation also describes capabilities such as:

  • network diagrams;
  • timelines;
  • lists;
  • cross-referencing;
  • sharing with users or access groups.

This matters because a company investigation often stops being a search task and becomes a case-management problem.

You may need to preserve:

  • your own candidate entities;
  • notes;
  • uploaded source documents;
  • relationship hypotheses;
  • cross-reference results.

OpenCorporates is primarily a company-data source.

Aleph can become part of the investigation environment itself.


Comparison at a glance

Research needOpenCorporatesOCCRP Aleph
establish exact legal company identitystrongest starting pointuseful when registry datasets are present
company number + jurisdiction anchorstrongcan store/use them
explicit provenance to public-source datastrong product emphasisdataset/document provenance depends on source
batch reconciliation of company namesstrongstrong cross-reference/reconciliation workflows
search across many heterogeneous datasetslimited compared with Alephcore strength
large document collectionsnot core rolecore strength
upload your own investigation materialnot core roleyes
network diagramsnot core roleyes
timelinesnot core roleyes
case/investigation workspacenoyes
company-registry discoverystrongdepends on datasets available
documentary contextlimitedstrong
best first uselegal entity anchorcontextual expansion / investigation

This table describes workflow fit.

It is not a universal quality score.


The strongest workflow is usually sequential

The tools are most useful when they are placed in the correct order.

A common sequence is:

messy company name
↓
OpenCorporates
↓
legal name + jurisdiction + company number + provenance
↓
primary registry verification
↓
Aleph entity
↓
cross-reference against documents and datasets
↓
investigate relevant matches
↓
preserve evidence

This sequence prevents one of the most common mistakes in corporate OSINT:

expanding a weak name match into a large investigation before identity is stable.


Worked example 1 — ambiguous company name

Suppose a document mentions:

Mercury Trading Ltd

You search broadly and find many similarly named entities.

OpenCorporates-first workflow

Use OpenCorporates to compare candidates by:

  • jurisdiction;
  • company number;
  • registered address;
  • status;
  • officers where available.

The document also contains an address.

One company matches:

Mercury Trading Ltd
Jurisdiction: GB
Company No: 87654321
Address: ...

Now you have a stronger entity anchor.

Move to primary source

Follow provenance or registry links and verify the important attributes against the relevant official company register.

Then move to Aleph

Create or locate the entity with:

  • legal name;
  • company number;
  • jurisdiction.

Cross-reference it across accessible datasets.

This order reduces false matches.


Worked example 2 — company already known, documentary context missing

You already know:

Example Infrastructure SA
Registration number: X
Jurisdiction: Y

The research question is:

Does this entity appear in public procurement or investigative document datasets?

This is not primarily an entity-discovery problem anymore.

Aleph may be the better first working environment.

Use the stable identifiers to:

  • search datasets;
  • create/cross-reference the entity;
  • inspect documents;
  • identify relevant relationships.

OpenCorporates can still be useful as an independent structured identity reference.

But it is no longer the center of the workflow.


Worked example 3 — list of 1,000 suppliers

You receive a spreadsheet:

supplier_name
country
address

You need to identify the underlying companies and then check whether any appear in investigative datasets.

Phase 1 — normalize legal entities

Use reconciliation against structured legal-entity data.

This is where OpenCorporates/OpenRefine-style reconciliation is useful.

Output should preserve:

input name
candidate legal entity
jurisdiction
company number
match decision
confidence/review state

Do not jump directly to contextual screening.

Phase 2 — cross-reference

Once the entity list is sufficiently clean, move those identifiers into Aleph-style cross-referencing.

Now the question changes:

Which normalized legal entities appear in other datasets?

This is exactly why entity resolution should precede contextual expansion.


Worked example 4 — uploaded leak or document collection

Suppose your investigation includes:

  • PDFs;
  • spreadsheets;
  • email archives.

The task is:

find which known companies and people appear in the material.

OpenCorporates is not a document-processing workspace.

Aleph is much closer to the problem.

You can:

  1. upload the documents to an investigation workspace;
  2. extract/search content;
  3. create or import entities;
  4. cross-reference entities;
  5. inspect matching documents;
  6. build network diagrams or timelines.

OpenCorporates can still help verify the legal companies you discover.

The tools now work in the opposite direction:

document discovery in Aleph
↓
candidate company
↓
OpenCorporates / primary registry
↓
legal identity verification

The sequence should follow the uncertainty.


Do not describe either tool too narrowly

A common comparison mistake is:

OpenCorporates = registries
Aleph = documents

That is too simplistic.

OpenCorporates can expose:

  • officers;
  • filings;
  • structured statements;
  • provenance;
  • other company-related public data.

Aleph can contain:

  • company registries;
  • sanctions lists;
  • structured entities;
  • many non-document datasets.

The more accurate distinction is:

OpenCorporates organizes its product around companies and legal-entity data.

Aleph organizes its product around heterogeneous investigative datasets, documents, entities and workspaces.

That is a difference in orientation.

Not a hard capability wall.


Provenance should follow every pivot

Whether you start in OpenCorporates or Aleph, preserve:

  • dataset or registry name;
  • source URL;
  • record identifier;
  • retrieval time;
  • entity identifiers;
  • document reference;
  • relationship type.

Do not allow a graph node to become detached from the source that created it.

A company record without provenance becomes hard to audit.

A network edge without provenance becomes a visual claim with no evidentiary foundation.


A company number is usually stronger than a name

Suppose Aleph shows:

Example Holdings Ltd

in ten datasets.

That looks impressive.

But if those datasets refer to several different companies with the same or similar name, the count is misleading.

A stronger investigation asks:

  • Which records contain the same company number?
  • Which jurisdiction?
  • Which address?
  • Which identifiers?
  • Which dates?

The goal is not to maximize matches.

It is to minimize identity ambiguity.


Dataset access affects Aleph results

Aleph's available results depend on the datasets and investigations a user has access to.

The platform supports access controls and group-based sharing.

This means:

no Aleph result

does not necessarily mean:

no relevant record exists anywhere.

It means:

no relevant result was found in the datasets available to that search under those access conditions.

That qualification matters.


OpenCorporates coverage also varies by jurisdiction and source

OpenCorporates relies on public company records.

Public-record availability differs across jurisdictions.

Some registries expose:

  • rich officer data;
  • filings;
  • addresses;
  • status history.

Others expose much less.

Therefore:

one OpenCorporates record can be more detailed than another because the underlying public sources differ.

Do not interpret a sparse record as proof that the company has a simple structure.

Coverage and corporate complexity are different things.


Freshness must remain visible

Corporate data changes.

Companies can:

  • change status;
  • appoint new officers;
  • move registered office;
  • file accounts;
  • change name;
  • merge;
  • dissolve.

Always record the retrieval time.

If OpenCorporates exposes provenance/retrieval metadata, preserve it.

If an Aleph dataset has a dataset date or source update context, preserve that too.

Avoid timeless writing such as:

Person A is director.

Prefer:

Registry-derived data retrieved on date X listed Person A as director.


Documentary mention is not legal identity

Aleph can surface a company name inside a document.

That is a documentary observation.

It does not automatically establish:

  • exact legal entity;
  • ownership;
  • involvement;
  • wrongdoing;
  • operational role.

The document needs interpretation.

Example:

Example Ltd appears in a spreadsheet.

Possible contexts:

  • customer;
  • supplier;
  • target;
  • counterparty;
  • historical mention;
  • unrelated copied data.

Resolve the entity and read the document context before assigning meaning.


Registry identity is not investigative significance

The reverse error also happens.

OpenCorporates may give you a clean legal company record.

That does not tell you:

  • why the company matters;
  • what a leaked record means;
  • whether it participated in the event you are researching.

Legal identity answers:

who is the entity?

Context answers:

why does the entity matter here?

Keep those questions separate.


When OpenCorporates should come first

Choose OpenCorporates first when:

  • you have only a company name;
  • jurisdiction is uncertain;
  • multiple similarly named entities exist;
  • you need company number;
  • you need a structured identity anchor;
  • you want provenance back to public company records;
  • you are reconciling a list of company names.

Typical first-stage output:

legal name
jurisdiction
company number
status
registered address
officers where available
provenance

Then decide whether broader context is needed.


When Aleph should come first

Choose Aleph first when:

  • the legal entity is already reasonably known;
  • the question is dataset/document-centric;
  • you need to search many heterogeneous sources;
  • you need to upload your own material;
  • you need cross-referencing;
  • you need network or timeline work;
  • you need an investigation workspace.

Typical first-stage output:

dataset matches
documents
related entities
candidate relationships
investigative leads

Then validate important entity identities and claims against stronger underlying records.


When you should use both

Use both when:

  • identity and context both matter;
  • a document mentions an ambiguous company;
  • a registry entity needs investigative expansion;
  • a large supplier/customer list must be normalized then cross-referenced;
  • an Aleph match needs legal verification;
  • an OpenCorporates entity needs documentary context.

This is probably the most common serious corporate-research pattern.


Decision matrix

SituationOpenCorporates firstAleph firstBoth
unknown company behind a nameyespossiblelater
exact company number knownusefulusefuloften
need official-source provenancestrongdepends on datasetyes
investigate uploaded PDFs/emailsnoyeslater for verification
batch legal-entity normalizationstrongusefulyes
cross-reference against many datasetslimitedstrongyes
build network diagramnostrongyes
verify Aleph company candidatestrongcandidate sourceyes
find documents mentioning known companylimitedstrongyes
case workspacenostrongno/yes as source

Again, this is workflow fit.

Not an overall score.


The best combined workflow

A robust company investigation can look like this.

1. Define the question

Example:

Does the company named in document X correspond to the entity registered in jurisdiction Y, and what other public records contextualize it?

2. Normalize the entity

Use OpenCorporates or another structured company source.

Preserve:

  • legal name;
  • company number;
  • jurisdiction;
  • provenance.

3. Verify the primary record

Check the official register when practical.

4. Create the investigation entity

Use the normalized identifiers in Aleph.

5. Cross-reference

Compare against available datasets.

6. Inspect actual records

Open the documents/dataset entries behind interesting matches.

7. Label the relationship

Do not leave graph edges semantically vague.

8. Verify important corporate facts

Return to registry/public-source evidence when necessary.

9. Preserve evidence

Record URLs, dataset IDs, document references and timestamps.

10. Write the conclusion

Separate:

  • legal identity;
  • documentary observation;
  • analytical inference.

Common mistakes

Mistake 1 — Asking which platform has "more data"

Volume is not the research question.

Mistake 2 — Searching Aleph before resolving a very ambiguous company name

Context expansion can multiply false candidates.

Mistake 3 — Treating OpenCorporates as the final authority without following provenance

Verify important records at the primary source when practical.

Mistake 4 — Treating an Aleph hit as proof of the exact legal entity

Inspect identifiers and document context.

Mistake 5 — Assuming Aleph contains every relevant dataset

Access and dataset availability matter.

Mistake 6 — Assuming sparse OpenCorporates data means the company has little activity

Source coverage varies by jurisdiction.

Mistake 7 — Treating a document mention as ownership or wrongdoing

The relationship must come from evidence.

Mistake 8 — Losing company numbers during the pivot

Identifiers are what make later matches defensible.

Mistake 9 — Building graphs without source references

Every important edge needs provenance.

Mistake 10 — Treating entity matching as binary too early

Keep candidate, verified, rejected and unresolved states explicit.


A compact rule for choosing

Ask what kind of uncertainty you have.

If the uncertainty is:

Which legal company is this?

start with OpenCorporates.

If the uncertainty is:

Where does this known entity appear across investigative data and documents?

start with Aleph.

If the uncertainty is:

Both

use them sequentially.

That simple rule resolves most tool-selection confusion.


Related OSINT.dev tools

OpenCorporates

Use as a structured company-data source for:

  • legal identity;
  • jurisdiction;
  • company numbers;
  • public-record provenance;
  • reconciliation.

OCCRP Aleph

Use for:

  • heterogeneous datasets;
  • documents;
  • cross-referencing;
  • investigations;
  • network diagrams;
  • timelines;
  • investigative context.

A Responsible Method for Company Research with Public Sources

Use the broader OSINT.dev guide when you need the full workflow around:

  • entity resolution;
  • official verification;
  • ownership/control;
  • sanctions screening;
  • document interpretation;
  • evidence preservation.

This comparison answers:

which of these two tools fits this stage?

The company-research guide answers:

how should the whole investigation be structured?


The core principle

OpenCorporates and Aleph are not competing answers to the same question.

They are strongest at different moments.

Think of the workflow this way:

OpenCorporates reduces legal-entity ambiguity.

Aleph increases investigative context around a sufficiently defined entity.

Then move back and forth as the investigation creates new uncertainty.

A document may create a new company candidate.

Return to structured entity resolution.

A registry record may reveal a new officer.

Move into cross-dataset context.

Good OSINT is iterative.

The important part is that every pivot retains:

  • identifiers;
  • provenance;
  • time;
  • relationship meaning;
  • uncertainty state.

The winner is not the tool with the larger result set.

The winner is the workflow that keeps the entity correct while the context expands.


References

Official product and user documentation used in this comparison:

tagsOSINTEthicalVerificationCompanyWorkflow
cite this article

OSINT.dev · Published Apr 21, 2026 · Updated Aug 20, 2026. Canonical URL: https://osint.dev/articles/opencorporates-vs-aleph-which-one-fits-which-research-job

03explore next

Related articles.

Editorial pieces that share a tool context or type with this one.