OpenCorporates vs OCCRP Aleph: Which One Fits Which Research Job?
OpenCorporates and OCCRP Aleph are both useful for company research.
They are also easy to misuse when they are treated as interchangeable.
Both can contain companies.
Both can help resolve identities.
Both can expose relationships.
Both can be used in larger investigative workflows.
But their center of gravity is different.
A useful starting distinction is:
OpenCorporates is strongest when you need a structured legal-entity anchor with explicit provenance.
Aleph is strongest when you need to search, organize and cross-reference entities and documents across multiple datasets inside an investigative workspace.
That distinction is not absolute.
OpenCorporates contains more than bare registry rows.
Aleph can contain company registries and can support entity reconciliation.
The point is not to force each platform into one box.
The point is to understand which research problem each one handles most naturally.
Start with the research question
Do not begin with:
Which tool is better?
Begin with:
What uncertainty am I trying to reduce?
Typical company-research questions include:
Identity questions
- Does this legal entity exist?
- What is its exact registered name?
- Which jurisdiction is it registered in?
- What is the company number?
- Is it active or dissolved?
- Which official source underlies the record?
These are usually entity-anchor questions.
Context questions
- Which documents mention this entity?
- Does it appear across multiple datasets?
- Which people, companies, contracts or assets are connected to it?
- Is it mentioned in leaks, procurement data, court records or sanctions lists?
- Can I compare a list of entities against a broader corpus?
These are usually context-expansion questions.
Investigation-management questions
- Can I upload my own documents?
- Can I build a case workspace?
- Can I draw relationships?
- Can I organize timelines?
- Can I cross-reference my own entities against available datasets?
These are investigation-workspace questions.
The right tool depends on which class of question you are asking now.
OpenCorporates: structured company identity first
OpenCorporates describes itself as a large open database of companies built from primary public sources.
Its API documentation emphasizes:
- structured company data;
- source and provenance information;
- company identifiers;
- jurisdiction;
- officers;
- filings and related public-record information where available;
- reconciliation of company names to legal corporate entities.
This makes it particularly strong when your first job is:
stabilize the legal entity before expanding the investigation.
Why provenance matters in OpenCorporates
A useful company record should not be a detached assertion.
You want to know:
- where the information came from;
- which publisher supplied it;
- when it was retrieved;
- whether it came from an external public source or another process.
OpenCorporates' API documentation exposes provenance/source information specifically for this reason.
That is analytically important.
Compare these two findings.
Weak:
OpenCorporates says Example Ltd is active.
Stronger:
OpenCorporates records Example Ltd, company number X in jurisdiction Y, and provides provenance to source Z, retrieved at time T.
The second statement preserves the path back to the underlying evidence.
For important findings, follow that provenance to the primary registry when practical.
OpenCorporates is especially useful for entity normalization
Company names are messy.
The same entity can appear as:
Example Holdings Limited
EXAMPLE HOLDINGS LTD
Example Holdings, Ltd.
Example Holdings
and different entities can share nearly identical names.
OpenCorporates provides reconciliation capabilities designed to help match company names to legal corporate entities.
Its OpenRefine reconciliation API is specifically intended for that use case.
This is useful when you start with:
- a spreadsheet of companies;
- names extracted from documents;
- supplier lists;
- counterparties;
- messy legacy data.
The objective is not:
find a similar-looking company.
It is:
identify the most plausible legal entity and retain the identifiers that let you distinguish it from others.
What OpenCorporates does not automatically solve
A structured legal-entity record does not answer every investigative question.
OpenCorporates does not automatically tell you:
- why an entity appears in a leaked document;
- whether a contract relationship is suspicious;
- what a court filing means;
- which documents in a large corpus mention the entity;
- how a person, company, asset and contract relate inside your specific investigation;
- whether a name match in another dataset is genuinely the same entity.
Even strong registry data still needs context.
That is where a platform like Aleph becomes useful.
OCCRP Aleph: multi-dataset and documentary context
Aleph is designed around datasets, documents and entities.
Its underlying FollowTheMoney model can represent many entity types, including:
- people;
- companies;
- assets;
- contracts;
- events;
- bank accounts;
- addresses;
- vessels and other structured objects.
Aleph datasets can include categories such as:
- company registries;
- sanctions lists;
- procurement;
- court archives;
- regulatory filings;
- leaks;
- document libraries;
- financial records;
- news archives.
That means Aleph is not merely a document search engine.
It is a system for bringing heterogeneous evidence into a common investigative model.
Cross-referencing is one of Aleph's defining strengths
Aleph's documentation describes cross-referencing as comparing entities from one dataset against entities in other datasets to find leads and patterns.
The critical detail is that good cross-referencing depends on keys that help distinguish one entity from another.
For a company, useful keys can include:
- registration number;
- jurisdiction;
- address;
- other identifiers;
- structured attributes.
This is the same principle used in disciplined entity resolution:
names generate candidates; identifiers reduce ambiguity.
Aleph therefore can contribute to entity resolution too.
The difference is that its main strength is often what happens after you have a useful entity representation:
compare it across other datasets and documents.
Aleph becomes more valuable as the investigation gets richer
Imagine you have already established:
Legal name: Example Holdings Ltd
Jurisdiction: GB
Company number: 12345678
Now you want to know:
- Does this entity appear in procurement datasets?
- Does it appear in a leak?
- Are documents connected to known officers?
- Do related entities appear elsewhere?
- Can I upload my own files and cross-reference them?
- Can I build a relationship graph or timeline?
This is where Aleph's model fits naturally.
The legal identity is the anchor.
Aleph helps turn the anchor into investigative context.
Aleph investigation workspaces are a different class of capability
Aleph supports investigation workspaces where users can upload and organize:
- entities;
- documents;
- email archives;
- spreadsheets and other files.
Its documentation also describes capabilities such as:
- network diagrams;
- timelines;
- lists;
- cross-referencing;
- sharing with users or access groups.
This matters because a company investigation often stops being a search task and becomes a case-management problem.
You may need to preserve:
- your own candidate entities;
- notes;
- uploaded source documents;
- relationship hypotheses;
- cross-reference results.
OpenCorporates is primarily a company-data source.
Aleph can become part of the investigation environment itself.
Comparison at a glance
| Research need | OpenCorporates | OCCRP Aleph |
|---|---|---|
| establish exact legal company identity | strongest starting point | useful when registry datasets are present |
| company number + jurisdiction anchor | strong | can store/use them |
| explicit provenance to public-source data | strong product emphasis | dataset/document provenance depends on source |
| batch reconciliation of company names | strong | strong cross-reference/reconciliation workflows |
| search across many heterogeneous datasets | limited compared with Aleph | core strength |
| large document collections | not core role | core strength |
| upload your own investigation material | not core role | yes |
| network diagrams | not core role | yes |
| timelines | not core role | yes |
| case/investigation workspace | no | yes |
| company-registry discovery | strong | depends on datasets available |
| documentary context | limited | strong |
| best first use | legal entity anchor | contextual expansion / investigation |
This table describes workflow fit.
It is not a universal quality score.
The strongest workflow is usually sequential
The tools are most useful when they are placed in the correct order.
A common sequence is:
messy company name
↓
OpenCorporates
↓
legal name + jurisdiction + company number + provenance
↓
primary registry verification
↓
Aleph entity
↓
cross-reference against documents and datasets
↓
investigate relevant matches
↓
preserve evidence
This sequence prevents one of the most common mistakes in corporate OSINT:
expanding a weak name match into a large investigation before identity is stable.
Worked example 1 — ambiguous company name
Suppose a document mentions:
Mercury Trading Ltd
You search broadly and find many similarly named entities.
OpenCorporates-first workflow
Use OpenCorporates to compare candidates by:
- jurisdiction;
- company number;
- registered address;
- status;
- officers where available.
The document also contains an address.
One company matches:
Mercury Trading Ltd
Jurisdiction: GB
Company No: 87654321
Address: ...
Now you have a stronger entity anchor.
Move to primary source
Follow provenance or registry links and verify the important attributes against the relevant official company register.
Then move to Aleph
Create or locate the entity with:
- legal name;
- company number;
- jurisdiction.
Cross-reference it across accessible datasets.
This order reduces false matches.
Worked example 2 — company already known, documentary context missing
You already know:
Example Infrastructure SA
Registration number: X
Jurisdiction: Y
The research question is:
Does this entity appear in public procurement or investigative document datasets?
This is not primarily an entity-discovery problem anymore.
Aleph may be the better first working environment.
Use the stable identifiers to:
- search datasets;
- create/cross-reference the entity;
- inspect documents;
- identify relevant relationships.
OpenCorporates can still be useful as an independent structured identity reference.
But it is no longer the center of the workflow.
Worked example 3 — list of 1,000 suppliers
You receive a spreadsheet:
supplier_name
country
address
You need to identify the underlying companies and then check whether any appear in investigative datasets.
Phase 1 — normalize legal entities
Use reconciliation against structured legal-entity data.
This is where OpenCorporates/OpenRefine-style reconciliation is useful.
Output should preserve:
input name
candidate legal entity
jurisdiction
company number
match decision
confidence/review state
Do not jump directly to contextual screening.
Phase 2 — cross-reference
Once the entity list is sufficiently clean, move those identifiers into Aleph-style cross-referencing.
Now the question changes:
Which normalized legal entities appear in other datasets?
This is exactly why entity resolution should precede contextual expansion.
Worked example 4 — uploaded leak or document collection
Suppose your investigation includes:
- PDFs;
- spreadsheets;
- email archives.
The task is:
find which known companies and people appear in the material.
OpenCorporates is not a document-processing workspace.
Aleph is much closer to the problem.
You can:
- upload the documents to an investigation workspace;
- extract/search content;
- create or import entities;
- cross-reference entities;
- inspect matching documents;
- build network diagrams or timelines.
OpenCorporates can still help verify the legal companies you discover.
The tools now work in the opposite direction:
document discovery in Aleph
↓
candidate company
↓
OpenCorporates / primary registry
↓
legal identity verification
The sequence should follow the uncertainty.
Do not describe either tool too narrowly
A common comparison mistake is:
OpenCorporates = registries
Aleph = documents
That is too simplistic.
OpenCorporates can expose:
- officers;
- filings;
- structured statements;
- provenance;
- other company-related public data.
Aleph can contain:
- company registries;
- sanctions lists;
- structured entities;
- many non-document datasets.
The more accurate distinction is:
OpenCorporates organizes its product around companies and legal-entity data.
Aleph organizes its product around heterogeneous investigative datasets, documents, entities and workspaces.
That is a difference in orientation.
Not a hard capability wall.
Provenance should follow every pivot
Whether you start in OpenCorporates or Aleph, preserve:
- dataset or registry name;
- source URL;
- record identifier;
- retrieval time;
- entity identifiers;
- document reference;
- relationship type.
Do not allow a graph node to become detached from the source that created it.
A company record without provenance becomes hard to audit.
A network edge without provenance becomes a visual claim with no evidentiary foundation.
A company number is usually stronger than a name
Suppose Aleph shows:
Example Holdings Ltd
in ten datasets.
That looks impressive.
But if those datasets refer to several different companies with the same or similar name, the count is misleading.
A stronger investigation asks:
- Which records contain the same company number?
- Which jurisdiction?
- Which address?
- Which identifiers?
- Which dates?
The goal is not to maximize matches.
It is to minimize identity ambiguity.
Dataset access affects Aleph results
Aleph's available results depend on the datasets and investigations a user has access to.
The platform supports access controls and group-based sharing.
This means:
no Aleph result
does not necessarily mean:
no relevant record exists anywhere.
It means:
no relevant result was found in the datasets available to that search under those access conditions.
That qualification matters.
OpenCorporates coverage also varies by jurisdiction and source
OpenCorporates relies on public company records.
Public-record availability differs across jurisdictions.
Some registries expose:
- rich officer data;
- filings;
- addresses;
- status history.
Others expose much less.
Therefore:
one OpenCorporates record can be more detailed than another because the underlying public sources differ.
Do not interpret a sparse record as proof that the company has a simple structure.
Coverage and corporate complexity are different things.
Freshness must remain visible
Corporate data changes.
Companies can:
- change status;
- appoint new officers;
- move registered office;
- file accounts;
- change name;
- merge;
- dissolve.
Always record the retrieval time.
If OpenCorporates exposes provenance/retrieval metadata, preserve it.
If an Aleph dataset has a dataset date or source update context, preserve that too.
Avoid timeless writing such as:
Person A is director.
Prefer:
Registry-derived data retrieved on date X listed Person A as director.
Documentary mention is not legal identity
Aleph can surface a company name inside a document.
That is a documentary observation.
It does not automatically establish:
- exact legal entity;
- ownership;
- involvement;
- wrongdoing;
- operational role.
The document needs interpretation.
Example:
Example Ltd appears in a spreadsheet.
Possible contexts:
- customer;
- supplier;
- target;
- counterparty;
- historical mention;
- unrelated copied data.
Resolve the entity and read the document context before assigning meaning.
Registry identity is not investigative significance
The reverse error also happens.
OpenCorporates may give you a clean legal company record.
That does not tell you:
- why the company matters;
- what a leaked record means;
- whether it participated in the event you are researching.
Legal identity answers:
who is the entity?
Context answers:
why does the entity matter here?
Keep those questions separate.
When OpenCorporates should come first
Choose OpenCorporates first when:
- you have only a company name;
- jurisdiction is uncertain;
- multiple similarly named entities exist;
- you need company number;
- you need a structured identity anchor;
- you want provenance back to public company records;
- you are reconciling a list of company names.
Typical first-stage output:
legal name
jurisdiction
company number
status
registered address
officers where available
provenance
Then decide whether broader context is needed.
When Aleph should come first
Choose Aleph first when:
- the legal entity is already reasonably known;
- the question is dataset/document-centric;
- you need to search many heterogeneous sources;
- you need to upload your own material;
- you need cross-referencing;
- you need network or timeline work;
- you need an investigation workspace.
Typical first-stage output:
dataset matches
documents
related entities
candidate relationships
investigative leads
Then validate important entity identities and claims against stronger underlying records.
When you should use both
Use both when:
- identity and context both matter;
- a document mentions an ambiguous company;
- a registry entity needs investigative expansion;
- a large supplier/customer list must be normalized then cross-referenced;
- an Aleph match needs legal verification;
- an OpenCorporates entity needs documentary context.
This is probably the most common serious corporate-research pattern.
Decision matrix
| Situation | OpenCorporates first | Aleph first | Both |
|---|---|---|---|
| unknown company behind a name | yes | possible | later |
| exact company number known | useful | useful | often |
| need official-source provenance | strong | depends on dataset | yes |
| investigate uploaded PDFs/emails | no | yes | later for verification |
| batch legal-entity normalization | strong | useful | yes |
| cross-reference against many datasets | limited | strong | yes |
| build network diagram | no | strong | yes |
| verify Aleph company candidate | strong | candidate source | yes |
| find documents mentioning known company | limited | strong | yes |
| case workspace | no | strong | no/yes as source |
Again, this is workflow fit.
Not an overall score.
The best combined workflow
A robust company investigation can look like this.
1. Define the question
Example:
Does the company named in document X correspond to the entity registered in jurisdiction Y, and what other public records contextualize it?
2. Normalize the entity
Use OpenCorporates or another structured company source.
Preserve:
- legal name;
- company number;
- jurisdiction;
- provenance.
3. Verify the primary record
Check the official register when practical.
4. Create the investigation entity
Use the normalized identifiers in Aleph.
5. Cross-reference
Compare against available datasets.
6. Inspect actual records
Open the documents/dataset entries behind interesting matches.
7. Label the relationship
Do not leave graph edges semantically vague.
8. Verify important corporate facts
Return to registry/public-source evidence when necessary.
9. Preserve evidence
Record URLs, dataset IDs, document references and timestamps.
10. Write the conclusion
Separate:
- legal identity;
- documentary observation;
- analytical inference.
Common mistakes
Mistake 1 — Asking which platform has "more data"
Volume is not the research question.
Mistake 2 — Searching Aleph before resolving a very ambiguous company name
Context expansion can multiply false candidates.
Mistake 3 — Treating OpenCorporates as the final authority without following provenance
Verify important records at the primary source when practical.
Mistake 4 — Treating an Aleph hit as proof of the exact legal entity
Inspect identifiers and document context.
Mistake 5 — Assuming Aleph contains every relevant dataset
Access and dataset availability matter.
Mistake 6 — Assuming sparse OpenCorporates data means the company has little activity
Source coverage varies by jurisdiction.
Mistake 7 — Treating a document mention as ownership or wrongdoing
The relationship must come from evidence.
Mistake 8 — Losing company numbers during the pivot
Identifiers are what make later matches defensible.
Mistake 9 — Building graphs without source references
Every important edge needs provenance.
Mistake 10 — Treating entity matching as binary too early
Keep candidate, verified, rejected and unresolved states explicit.
A compact rule for choosing
Ask what kind of uncertainty you have.
If the uncertainty is:
Which legal company is this?
start with OpenCorporates.
If the uncertainty is:
Where does this known entity appear across investigative data and documents?
start with Aleph.
If the uncertainty is:
Both
use them sequentially.
That simple rule resolves most tool-selection confusion.
Related OSINT.dev tools
OpenCorporates
Use as a structured company-data source for:
- legal identity;
- jurisdiction;
- company numbers;
- public-record provenance;
- reconciliation.
OCCRP Aleph
Use for:
- heterogeneous datasets;
- documents;
- cross-referencing;
- investigations;
- network diagrams;
- timelines;
- investigative context.
A Responsible Method for Company Research with Public Sources
Use the broader OSINT.dev guide when you need the full workflow around:
- entity resolution;
- official verification;
- ownership/control;
- sanctions screening;
- document interpretation;
- evidence preservation.
This comparison answers:
which of these two tools fits this stage?
The company-research guide answers:
how should the whole investigation be structured?
The core principle
OpenCorporates and Aleph are not competing answers to the same question.
They are strongest at different moments.
Think of the workflow this way:
OpenCorporates reduces legal-entity ambiguity.
Aleph increases investigative context around a sufficiently defined entity.
Then move back and forth as the investigation creates new uncertainty.
A document may create a new company candidate.
Return to structured entity resolution.
A registry record may reveal a new officer.
Move into cross-dataset context.
Good OSINT is iterative.
The important part is that every pivot retains:
- identifiers;
- provenance;
- time;
- relationship meaning;
- uncertainty state.
The winner is not the tool with the larger result set.
The winner is the workflow that keeps the entity correct while the context expands.
References
Official product and user documentation used in this comparison:
-
OpenCorporates API — official API and data overview
https://api.opencorporates.com/ -
OpenCorporates API Reference — provenance, company data and source model
https://api.opencorporates.com/documentation/API-Reference -
OpenCorporates OpenRefine Reconciliation API — matching company names to legal entities
https://api.opencorporates.com/documentation/Open-Refine-Reconciliation-API -
OCCRP Aleph — User and Technical Documentation
https://docs.aleph.occrp.org/ -
OCCRP Aleph — Key Terms and FollowTheMoney entities
https://docs.aleph.occrp.org/users/getting-started/key-terms/ -
OCCRP Aleph — Cross-referencing your data
https://docs.aleph.occrp.org/users/investigations/cross-referencing/ -
OCCRP Aleph — Investigation workspaces
https://docs.aleph.occrp.org/users/investigations/overview/ -
OCCRP Aleph — Network diagrams
https://docs.aleph.occrp.org/users/investigations/network-diagrams/
OSINT.dev · Published Apr 21, 2026 · Updated Aug 20, 2026. Canonical URL: https://osint.dev/articles/opencorporates-vs-aleph-which-one-fits-which-research-job
Related articles.
Editorial pieces that share a tool context or type with this one.
How to Use Sanctions and Risk Lists Without Overreading Them
A disciplined method for sanctions, PEP and risk-list research: resolve the entity, classify the signal, verify the authority, distinguish ownership and control rules, preserve time, and avoid turning adjacency into a verdict.
Start Here: How to Use an OSINT Tool Catalog Without Getting Lost
A practical starting guide to choosing OSINT tools by question, signal family, risk and evidence needs — and using the catalog as a decision system instead of a link directory.
A Responsible Method for Company Research with Public Sources
A practical framework for resolving legal entities, verifying registry records, mapping ownership and control, screening sanctions carefully, and preserving evidence without turning name matches into conclusions.
Stop Calling Them Hackers
Calling every cyber actor a hacker collapses authorization, motive and attribution into one vague label. Better OSINT starts with language precise enough to preserve what the evidence actually says.