guide · featured

A Responsible Method for Company Research with Public Sources

A practical framework for resolving legal entities, verifying registry records, mapping ownership and control, screening sanctions carefully, and preserving evidence without turning name matches into conclusions.

published
Apr 21, 2026
updated
Aug 18, 2026
slug
a-responsible-method-for-company-research-with-public-sources
status
Published

A Responsible Method for Company Research with Public Sources

Company research looks simple until two companies share the same name.

Or one company has changed its legal name.

Or a director appears in several entities.

Or a sanctions dataset returns a partial match.

Or a document mentions a company without proving that the company participated in the activity described.

At that point, the task is no longer:

Search for a company.

It becomes:

Resolve the legal entity, identify the strongest public records, reconstruct relevant relationships, screen risk signals carefully, preserve provenance, and write conclusions no stronger than the evidence supports.

That sequence matters.

The most common company-research errors happen when analysts reverse it:

find an interesting risk signal → search for a matching name → assume identity → build a story around the match.

A responsible workflow starts with identity.

Only then does it move to relationships, documents and risk.


The first problem is entity resolution

Company names are not reliable unique identifiers.

Consider:

Atlas Consulting Ltd

There may be:

  • multiple companies with similar names;
  • companies in different jurisdictions;
  • dissolved and active entities;
  • branches;
  • historical names;
  • trade names;
  • subsidiaries;
  • parent companies.

So the first question should not be:

What can I find about Atlas Consulting?

It should be:

Which legal entity am I researching?

A minimum identity record

Before deeper research, try to establish:

  • official legal name;
  • jurisdiction;
  • company or registry number;
  • legal form;
  • incorporation or registration date;
  • current legal status;
  • registered address;
  • known previous names;
  • authoritative source URL where available;
  • other persistent identifiers such as LEI where relevant.

The strongest identity record combines several attributes.

For example:

Legal name: Example Holdings S.r.l.
Jurisdiction: Italy
Registry number: ...
Registered address: ...
Status: Active

is far more useful than:

Name: Example Holdings

Prefer identifiers over names

Names are useful for discovery.

Identifiers are better for resolution.

Depending on jurisdiction and context, useful identifiers can include:

  • company registration numbers;
  • tax or business identifiers where public;
  • Legal Entity Identifiers;
  • regulator identifiers;
  • securities identifiers;
  • official filing numbers.

The Global Legal Entity Identifier Foundation describes the LEI as a unique identifier for legal entities participating in financial transactions and other official interactions.

GLEIF reference data can contain information such as:

  • official legal name;
  • registered address;
  • entity status;
  • registration authority information;
  • relationship data where available.

Not every company has an LEI.

The absence of an LEI is therefore not evidence that a company does not exist.

Use the identifier when it exists and fits the context.


Source hierarchy matters

Not every company database has the same evidentiary role.

A useful hierarchy is:

Level 1 — Official registry or authority

Examples include:

  • national company registers;
  • securities regulators;
  • financial regulators;
  • procurement authorities;
  • court or insolvency registers;
  • official gazettes.

When accessible and current, these are often the strongest sources for legal facts such as:

  • registration;
  • status;
  • filings;
  • named officers;
  • official addresses.

Even official records require interpretation.

A registry tells you what was filed or recorded.

It does not automatically prove that every filing describes current operational reality.

Level 2 — Structured aggregators with provenance

OpenCorporates is an example.

Its company records are based on public records published by company registers and other public sources, and its data model can include source links and provenance.

This can be extremely useful for cross-jurisdiction discovery and normalization.

But the safe workflow is:

aggregator → identify source → verify important facts against the underlying official record when possible.

Do not turn:

OpenCorporates lists X

into:

the official register currently confirms X

unless you have actually checked the official record.

Level 3 — Investigative datasets and document platforms

OCCRP Aleph belongs here.

Aleph can contain:

  • public datasets;
  • structured entities;
  • documents;
  • leaks;
  • filings;
  • investigation workspaces;
  • relationships.

It is especially useful for cross-referencing names and identifiers across many datasets.

But a match inside an investigative corpus is not automatically a legal fact.

You still need to ask:

  • Which dataset did the result come from?
  • What is the source document?
  • What date does it describe?
  • Is the entity identity resolved?
  • What exactly does the document assert?

Level 4 — Search, media and first-party claims

Examples include:

  • company websites;
  • press releases;
  • news articles;
  • social profiles;
  • search results.

These can provide valuable context.

They are not substitutes for entity resolution.


Step 1 — Define the research question

Company research expands quickly.

Start with a narrow question.

Examples:

Does this company legally exist in the jurisdiction it claims?

Who are the currently listed officers?

Is company A legally connected to company B?

Does the entity appear in a sanctions dataset?

Has the company changed name or status?

Is a website's claimed legal entity consistent with public records?

Each question requires different evidence.

If your question is:

Does the company exist?

you do not yet need a sanctions graph.

If your question is:

Who controls this company?

the legal existence record is necessary but not sufficient.


Step 2 — Establish the canonical entity

Create a compact identity card.

Example:

Research entity
---------------
Legal name:
Jurisdiction:
Registration number:
Legal form:
Status:
Registered address:
Incorporation date:
Previous names:
Primary registry source:
Other identifiers:
Last checked:

Do not proceed to high-impact conclusions until this identity is reasonably stable.

Handle ambiguity explicitly

If two possible entities remain, preserve both candidates.

For example:

Candidate A
Example Ltd
UK company number 12345678

Candidate B
Example Limited
Irish CRO number 987654

Then collect discriminators:

  • address;
  • officer;
  • website;
  • incorporation date;
  • country;
  • LEI;
  • filing references.

A good analyst is comfortable writing:

Identity unresolved between two plausible entities.

That is stronger than forcing a match.


Step 3 — Verify status and time

Company data is time-sensitive.

A company can be:

  • active;
  • dissolved;
  • struck off;
  • in liquidation;
  • merged;
  • renamed;
  • converted;
  • reinstated.

Always attach a date to status.

Prefer:

The registry showed the entity as active on 18 August 2026.

over:

The company is active.

The first statement preserves the temporal evidence.

Historical names matter

A previous legal name can explain:

  • old contracts;
  • old websites;
  • historical sanctions references;
  • archived documents;
  • previous domain registrations;
  • older media coverage.

Do not treat an old name as a different entity without checking identifiers.

Likewise, do not assume identical names mean the same entity.


Step 4 — Map people and related companies carefully

Once the entity is stable, begin relationship research.

Possible public relationships include:

  • directors;
  • officers;
  • beneficial owners where legally public;
  • shareholders where public;
  • parent/subsidiary relationships;
  • branches;
  • registered agents;
  • joint ventures;
  • counterparties.

Every relationship should contain:

Entity A
Relation type
Entity B
Source
Date or period
Confidence
Notes

This prevents graphs from becoming anonymous lines.

A graph edge is a claim

A network diagram can make a weak relationship look strong.

For example:

Person A → Company B

could mean:

  • director;
  • former director;
  • shareholder;
  • beneficial owner;
  • document mention;
  • address co-occurrence;
  • unknown inferred association.

Those are not equivalent.

Always label the relationship.


Aleph: cross-reference, then inspect provenance

OCCRP Aleph is useful when the investigation moves from one legal entity to a larger document or relationship context.

Aleph's cross-referencing workflow uses identifying keys that help distinguish entities.

That is a crucial concept.

Do not cross-reference only:

John Smith

when you can also provide:

  • date of birth;
  • nationality;
  • company number;
  • address;
  • identifier.

The stronger the entity representation, the more meaningful the candidate matches.

A match is not a conclusion

If Aleph returns a company in a dataset, ask:

  1. Which dataset?
  2. Which record?
  3. Which identifying fields matched?
  4. Is the match exact, approximate or inferred?
  5. What is the underlying document?
  6. What does that document actually say?

Then record the finding at the appropriate strength.


Step 5 — Separate ownership from control

People often use these terms interchangeably.

They should not.

Ownership

Ownership generally refers to an economic or legal stake.

Public data may describe:

  • direct ownership;
  • indirect ownership;
  • percentage;
  • share class;
  • ownership period.

Control

Control may arise through:

  • voting rights;
  • board powers;
  • agreements;
  • management roles;
  • parent entities;
  • other legal mechanisms.

A director is not automatically an owner.

An owner is not automatically an operational controller.

A parent relationship does not automatically explain every subsidiary action.

If a source does not specify the relation, do not invent it.


Beneficial ownership requires special care

Beneficial ownership information can be powerful.

It is also jurisdiction-dependent and frequently incomplete.

Availability may depend on:

  • local law;
  • legitimate-interest rules;
  • regulated access;
  • filing thresholds;
  • disclosure exemptions;
  • source freshness.

When beneficial ownership data is available, record:

  • exact source;
  • legal definition used;
  • percentage or control basis if stated;
  • effective date;
  • whether the information is declared, verified or derived.

Do not generalize one jurisdiction's ownership definition to another.


Step 6 — Use sanctions and risk data after identity resolution

Risk screening should come after entity resolution.

OpenSanctions provides data and matching workflows for sanctions, politically exposed persons and other risk topics.

Its matching API returns candidate matches based on the properties supplied.

That wording is important.

A screening result is not:

This company is sanctioned.

It is:

The supplied entity produced candidate match X with these identifying attributes and this source context.

Then you verify.


Name-only sanctions matching is weak

Suppose your research company is:

Global Trading Ltd

A risk database also contains:

Global Trading Limited

That is not enough.

Compare:

  • jurisdiction;
  • registration number;
  • address;
  • incorporation date;
  • identifiers;
  • directors;
  • nationality where relevant;
  • source list.

A name similarity can be a useful lead.

It is not identity proof.


Direct listing and adjacency are different findings

These statements have very different meanings:

Company X is directly designated on sanctions list Y.

Company X is owned by a designated entity.

A director of Company X is designated.

Company X appears in a dataset that also contains sanctioned entities.

Company X has a commercial relationship with an entity that appears on a risk list.

Do not compress all of them into:

Company X is sanctioned.

The relation type is part of the finding.


PEP is not a synonym for criminal or sanctioned

A politically exposed person classification is a risk-management category used in financial-crime compliance contexts.

It is not an accusation of wrongdoing.

Likewise:

  • relative or close associate;
  • state-owned enterprise;
  • sanctioned entity;
  • sanctions-linked entity;

are distinct concepts.

If your source uses a semantic risk topic, preserve the exact topic.

Avoid moral language the data does not support.


Trace risk data back to the underlying source

Aggregators are valuable because they normalize many lists.

For important findings, inspect:

  • source dataset;
  • issuing authority;
  • program;
  • designation date;
  • identifiers;
  • current status;
  • any removal or amendment history.

The strongest statement is tied to the authority.

For example:

OpenSanctions aggregates record X from authority Y.

Then verify authority Y when practical.


Step 7 — Use documents to answer relationship questions

Corporate research frequently turns into document research.

Useful documents can include:

  • annual accounts;
  • incorporation documents;
  • shareholder filings;
  • official notices;
  • procurement records;
  • court filings;
  • regulatory actions;
  • contracts;
  • leaked documents.

Documents need provenance.

Record:

Document title
Source
Document date
Publication date
Pages
Entity identifiers
Relevant passage
Interpretation

A company name appearing in a document proves only that the document contains that name.

The significance depends on context.


Document mention is not participation

Suppose a leaked spreadsheet contains:

Example Holdings Ltd

Possible interpretations include:

  • customer;
  • supplier;
  • owner;
  • target;
  • unrelated reference;
  • copied address book;
  • historical record.

The document itself must establish the relation.

Avoid:

Company X was involved.

when the evidence only supports:

Company X is mentioned in document Y.

This is one of the most important writing disciplines in corporate OSINT.


Step 8 — Connect the company's claimed web identity

A company's website can help corroborate legal identity.

Useful first-party signals include:

  • legal footer;
  • registered address;
  • company number;
  • privacy notice;
  • terms;
  • contact page;
  • corporate group statement.

Compare those against registry data.

Strong consistency

Website legal name = registry legal name
Company number matches
Address matches

This supports the relationship between the website and legal entity.

Weak consistency

Brand name only
No registration number
Generic contact form

The website may still be legitimate, but the entity linkage is weaker.

Contradiction

Website claims Company A
Terms identify Company B
Registry address belongs to Company C

That deserves investigation.

Do not automatically assume fraud.

Corporate groups, trading names and outsourced services can create legitimate complexity.


Historical websites can explain corporate change

The Wayback Machine can help answer:

  • Did the site previously identify another legal entity?
  • Did the registered address change?
  • Did branding change around a merger?
  • Did privacy terms name a previous company?
  • Did a domain previously belong to another business?

Historical web evidence is especially useful when legal records show:

  • name change;
  • acquisition;
  • dissolution;
  • restructuring.

Archive evidence remains incomplete.

A missing capture does not prove absence.


Step 9 — Preserve evidence before the research expands

Company research produces many volatile sources.

Some can change:

  • registry interfaces;
  • websites;
  • sanctions entries;
  • press releases;
  • online filings.

For important findings, preserve:

  • URL;
  • source name;
  • retrieval time;
  • identifiers;
  • relevant text or record;
  • local copy or capture where appropriate and lawful.

Tools such as Hunchly can help package investigative browsing and evidence trails.

Public archives can preserve web states.

These are different preservation models.

Choose the one appropriate to the investigation.


Step 10 — Build a structured evidence ledger

A useful company research ledger might contain:

ClaimEvidenceSourceDateConfidenceNotes
Company legally existsregistry recordofficial registrydatehighnumber X
Website linked to entitylegal footer + company numbercompany websitedatehighmatches registry
Person A is directorfilingregistrydatehighcurrent as of date
Person A owns companyunknownunresolveddirector ≠ owner
Sanctions candidate existsOpenSanctions candidateaggregated sourcedatemediumidentity review required

The ledger prevents weak signals from silently becoming facts.


Use confidence deliberately

A simple confidence model can help.

High confidence

Use when:

  • authoritative identifier matches;
  • official record is clear;
  • multiple independent sources agree;
  • relation type is explicit.

Medium confidence

Use when:

  • evidence is strong but incomplete;
  • structured aggregator and first-party data agree;
  • document context supports the relation but one element remains unresolved.

Low confidence

Use when:

  • name-only match;
  • single unverified secondary source;
  • weak address similarity;
  • inferred relationship;
  • stale or ambiguous record.

Do not hide low confidence.

Low-confidence findings can still guide the next step.


A worked example

Suppose a website claims to belong to:

Northstar Logistics

You need to determine the underlying company and check whether a sanctions-related claim circulating online is accurate.

Phase 1 — Resolve the company

The website footer says:

Northstar Logistics Ltd
Company No. 12345678

OpenCorporates returns a company with:

  • matching name;
  • matching jurisdiction;
  • matching company number.

You follow the provenance to the official registry.

The registry confirms:

  • company number;
  • active status;
  • registered office.

Finding

The website's stated legal entity is consistent with the official company registered under number 12345678 as of the checked date.

This is a strong identity finding.


Phase 2 — Officers

The registry names:

Person A

as a current director.

Record:

Person A is listed as a director.

Do not write:

Person A owns the company.

unless ownership evidence supports that.


Phase 3 — Sanctions screening

A search for:

Northstar Logistics

returns a similarly named sanctions candidate in another jurisdiction.

The candidate has:

  • different registration number;
  • different address;
  • different country.

Finding

A similarly named sanctioned entity exists, but the available identifiers do not support treating it as the same legal entity.

This negative resolution is valuable.

You prevented a false positive.


Phase 4 — Investigative documents

Aleph contains a document mentioning:

Northstar Logistics Ltd

with the correct company number.

The document places the entity in a supplier list.

Finding

The entity is identified as a supplier in document Y.

Do not transform:

supplier

into:

owner, conspirator or controlled company

without additional evidence.


Phase 5 — Historical website

A Wayback capture shows the same domain previously listing a different legal entity before a documented acquisition.

Now the timeline has context.

Calibrated conclusion

A defensible summary could be:

Northstar Logistics Ltd, registration number 12345678, is the legal entity currently identified by the website and registry. Person A is listed as a director. A sanctions search produced a similarly named entity, but the available identifiers do not support identity. Historical web evidence indicates that the domain previously identified a different legal entity before the current corporate structure.

Every sentence has a different evidence basis.

That is what good company research should look like.


Common mistakes

Mistake 1 — Starting with the company name instead of the entity

Names collide.

Use identifiers.

Mistake 2 — Treating OpenCorporates as the registry itself

It is a structured source that can provide provenance and links.

Verify important facts against the underlying register when practical.

Mistake 3 — Treating every Aleph match as identity

Cross-reference results need entity-resolution review.

Mistake 4 — Treating a sanctions candidate as a confirmed match

Screening is matching.

Matching produces candidates.

Mistake 5 — Calling a PEP "sanctioned"

These are different categories.

Mistake 6 — Confusing director, owner and controller

Record the exact relationship.

Mistake 7 — Treating document mention as participation

A mention requires context.

Mistake 8 — Ignoring dates

Corporate structures change.

Mistake 9 — Building an unlabeled network graph

Every edge needs a source and relation type.

Mistake 10 — Researching every connected entity

Expand only when a relationship is relevant to the original question.


A repeatable company-research workflow

Use this sequence.

1. Define the question

What exactly are you trying to establish?

2. Resolve the entity

Collect:

  • legal name;
  • jurisdiction;
  • registry number;
  • identifiers.

3. Verify official status

Check the strongest available authoritative record.

4. Record names and identifiers

Preserve:

  • current name;
  • previous names;
  • company numbers;
  • LEI where relevant.

5. Map explicit relationships

Directors, owners, parents, subsidiaries or other relationships only when supported.

6. Search document and investigative sources

Use entity keys, not names alone.

7. Screen sanctions and risk data

Only after identity is stable.

8. Review the underlying risk source

Determine:

  • direct listing;
  • relationship;
  • topic;
  • date;
  • source authority.

9. Corroborate the company's web identity

Compare first-party claims to legal records.

10. Add historical context

Use archival and dated records when the question requires a timeline.

11. Preserve evidence

Keep provenance before interfaces or pages change.

12. Write facts, inferences and unresolved questions separately

Then stop when the research question has been answered.


Related OSINT.dev tools

OpenCorporates

Use for cross-jurisdiction company discovery and structured company records.

Treat provenance and official-source links as part of the result.

OCCRP Aleph

Use when the research moves into:

  • large document collections;
  • cross-referencing;
  • entity relationships;
  • investigative datasets.

Resolve entity identity before interpreting matches.

OpenSanctions

Use for sanctions, PEP and related risk-data screening.

A matching result should be treated as a candidate that requires identity validation and source review.

Wayback Machine

Use historical public pages to compare:

  • legal footers;
  • company names;
  • group ownership statements;
  • addresses;
  • branding.

Hunchly

Use case-oriented capture when the investigation requires a reproducible browsing and evidence trail.


The core principle

Responsible company research is not a search for suspicious connections.

It is a process of identity resolution and evidence control.

Use this order:

question → legal identity → authoritative record → explicit relationships → documents → risk screening → corroboration → preservation → calibrated conclusion

The order protects the investigation from its most dangerous error:

finding a compelling story before establishing that the records refer to the same entity.

A name is a lead.

An identifier is stronger.

A relationship needs a source.

A risk match needs validation.

A document needs context.

And every conclusion should be traceable back to the evidence that supports it.


References

Primary and official product/data documentation used in this guide:

tagsOSINTEthicalVerificationCompany
cite this article

OSINT.dev · Published Apr 21, 2026 · Updated Aug 18, 2026. Canonical URL: https://osint.dev/articles/a-responsible-method-for-company-research-with-public-sources

03explore next

Related articles.

Editorial pieces that share a tool context or type with this one.