A Responsible Method for Company Research with Public Sources
Company research looks simple until two companies share the same name.
Or one company has changed its legal name.
Or a director appears in several entities.
Or a sanctions dataset returns a partial match.
Or a document mentions a company without proving that the company participated in the activity described.
At that point, the task is no longer:
Search for a company.
It becomes:
Resolve the legal entity, identify the strongest public records, reconstruct relevant relationships, screen risk signals carefully, preserve provenance, and write conclusions no stronger than the evidence supports.
That sequence matters.
The most common company-research errors happen when analysts reverse it:
find an interesting risk signal → search for a matching name → assume identity → build a story around the match.
A responsible workflow starts with identity.
Only then does it move to relationships, documents and risk.
The first problem is entity resolution
Company names are not reliable unique identifiers.
Consider:
Atlas Consulting Ltd
There may be:
- multiple companies with similar names;
- companies in different jurisdictions;
- dissolved and active entities;
- branches;
- historical names;
- trade names;
- subsidiaries;
- parent companies.
So the first question should not be:
What can I find about Atlas Consulting?
It should be:
Which legal entity am I researching?
A minimum identity record
Before deeper research, try to establish:
- official legal name;
- jurisdiction;
- company or registry number;
- legal form;
- incorporation or registration date;
- current legal status;
- registered address;
- known previous names;
- authoritative source URL where available;
- other persistent identifiers such as LEI where relevant.
The strongest identity record combines several attributes.
For example:
Legal name: Example Holdings S.r.l.
Jurisdiction: Italy
Registry number: ...
Registered address: ...
Status: Active
is far more useful than:
Name: Example Holdings
Prefer identifiers over names
Names are useful for discovery.
Identifiers are better for resolution.
Depending on jurisdiction and context, useful identifiers can include:
- company registration numbers;
- tax or business identifiers where public;
- Legal Entity Identifiers;
- regulator identifiers;
- securities identifiers;
- official filing numbers.
The Global Legal Entity Identifier Foundation describes the LEI as a unique identifier for legal entities participating in financial transactions and other official interactions.
GLEIF reference data can contain information such as:
- official legal name;
- registered address;
- entity status;
- registration authority information;
- relationship data where available.
Not every company has an LEI.
The absence of an LEI is therefore not evidence that a company does not exist.
Use the identifier when it exists and fits the context.
Source hierarchy matters
Not every company database has the same evidentiary role.
A useful hierarchy is:
Level 1 — Official registry or authority
Examples include:
- national company registers;
- securities regulators;
- financial regulators;
- procurement authorities;
- court or insolvency registers;
- official gazettes.
When accessible and current, these are often the strongest sources for legal facts such as:
- registration;
- status;
- filings;
- named officers;
- official addresses.
Even official records require interpretation.
A registry tells you what was filed or recorded.
It does not automatically prove that every filing describes current operational reality.
Level 2 — Structured aggregators with provenance
OpenCorporates is an example.
Its company records are based on public records published by company registers and other public sources, and its data model can include source links and provenance.
This can be extremely useful for cross-jurisdiction discovery and normalization.
But the safe workflow is:
aggregator → identify source → verify important facts against the underlying official record when possible.
Do not turn:
OpenCorporates lists X
into:
the official register currently confirms X
unless you have actually checked the official record.
Level 3 — Investigative datasets and document platforms
OCCRP Aleph belongs here.
Aleph can contain:
- public datasets;
- structured entities;
- documents;
- leaks;
- filings;
- investigation workspaces;
- relationships.
It is especially useful for cross-referencing names and identifiers across many datasets.
But a match inside an investigative corpus is not automatically a legal fact.
You still need to ask:
- Which dataset did the result come from?
- What is the source document?
- What date does it describe?
- Is the entity identity resolved?
- What exactly does the document assert?
Level 4 — Search, media and first-party claims
Examples include:
- company websites;
- press releases;
- news articles;
- social profiles;
- search results.
These can provide valuable context.
They are not substitutes for entity resolution.
Step 1 — Define the research question
Company research expands quickly.
Start with a narrow question.
Examples:
Does this company legally exist in the jurisdiction it claims?
Who are the currently listed officers?
Is company A legally connected to company B?
Does the entity appear in a sanctions dataset?
Has the company changed name or status?
Is a website's claimed legal entity consistent with public records?
Each question requires different evidence.
If your question is:
Does the company exist?
you do not yet need a sanctions graph.
If your question is:
Who controls this company?
the legal existence record is necessary but not sufficient.
Step 2 — Establish the canonical entity
Create a compact identity card.
Example:
Research entity
---------------
Legal name:
Jurisdiction:
Registration number:
Legal form:
Status:
Registered address:
Incorporation date:
Previous names:
Primary registry source:
Other identifiers:
Last checked:
Do not proceed to high-impact conclusions until this identity is reasonably stable.
Handle ambiguity explicitly
If two possible entities remain, preserve both candidates.
For example:
Candidate A
Example Ltd
UK company number 12345678
Candidate B
Example Limited
Irish CRO number 987654
Then collect discriminators:
- address;
- officer;
- website;
- incorporation date;
- country;
- LEI;
- filing references.
A good analyst is comfortable writing:
Identity unresolved between two plausible entities.
That is stronger than forcing a match.
Step 3 — Verify status and time
Company data is time-sensitive.
A company can be:
- active;
- dissolved;
- struck off;
- in liquidation;
- merged;
- renamed;
- converted;
- reinstated.
Always attach a date to status.
Prefer:
The registry showed the entity as active on 18 August 2026.
over:
The company is active.
The first statement preserves the temporal evidence.
Historical names matter
A previous legal name can explain:
- old contracts;
- old websites;
- historical sanctions references;
- archived documents;
- previous domain registrations;
- older media coverage.
Do not treat an old name as a different entity without checking identifiers.
Likewise, do not assume identical names mean the same entity.
Step 4 — Map people and related companies carefully
Once the entity is stable, begin relationship research.
Possible public relationships include:
- directors;
- officers;
- beneficial owners where legally public;
- shareholders where public;
- parent/subsidiary relationships;
- branches;
- registered agents;
- joint ventures;
- counterparties.
Every relationship should contain:
Entity A
Relation type
Entity B
Source
Date or period
Confidence
Notes
This prevents graphs from becoming anonymous lines.
A graph edge is a claim
A network diagram can make a weak relationship look strong.
For example:
Person A → Company B
could mean:
- director;
- former director;
- shareholder;
- beneficial owner;
- document mention;
- address co-occurrence;
- unknown inferred association.
Those are not equivalent.
Always label the relationship.
Aleph: cross-reference, then inspect provenance
OCCRP Aleph is useful when the investigation moves from one legal entity to a larger document or relationship context.
Aleph's cross-referencing workflow uses identifying keys that help distinguish entities.
That is a crucial concept.
Do not cross-reference only:
John Smith
when you can also provide:
- date of birth;
- nationality;
- company number;
- address;
- identifier.
The stronger the entity representation, the more meaningful the candidate matches.
A match is not a conclusion
If Aleph returns a company in a dataset, ask:
- Which dataset?
- Which record?
- Which identifying fields matched?
- Is the match exact, approximate or inferred?
- What is the underlying document?
- What does that document actually say?
Then record the finding at the appropriate strength.
Step 5 — Separate ownership from control
People often use these terms interchangeably.
They should not.
Ownership
Ownership generally refers to an economic or legal stake.
Public data may describe:
- direct ownership;
- indirect ownership;
- percentage;
- share class;
- ownership period.
Control
Control may arise through:
- voting rights;
- board powers;
- agreements;
- management roles;
- parent entities;
- other legal mechanisms.
A director is not automatically an owner.
An owner is not automatically an operational controller.
A parent relationship does not automatically explain every subsidiary action.
If a source does not specify the relation, do not invent it.
Beneficial ownership requires special care
Beneficial ownership information can be powerful.
It is also jurisdiction-dependent and frequently incomplete.
Availability may depend on:
- local law;
- legitimate-interest rules;
- regulated access;
- filing thresholds;
- disclosure exemptions;
- source freshness.
When beneficial ownership data is available, record:
- exact source;
- legal definition used;
- percentage or control basis if stated;
- effective date;
- whether the information is declared, verified or derived.
Do not generalize one jurisdiction's ownership definition to another.
Step 6 — Use sanctions and risk data after identity resolution
Risk screening should come after entity resolution.
OpenSanctions provides data and matching workflows for sanctions, politically exposed persons and other risk topics.
Its matching API returns candidate matches based on the properties supplied.
That wording is important.
A screening result is not:
This company is sanctioned.
It is:
The supplied entity produced candidate match X with these identifying attributes and this source context.
Then you verify.
Name-only sanctions matching is weak
Suppose your research company is:
Global Trading Ltd
A risk database also contains:
Global Trading Limited
That is not enough.
Compare:
- jurisdiction;
- registration number;
- address;
- incorporation date;
- identifiers;
- directors;
- nationality where relevant;
- source list.
A name similarity can be a useful lead.
It is not identity proof.
Direct listing and adjacency are different findings
These statements have very different meanings:
Company X is directly designated on sanctions list Y.
Company X is owned by a designated entity.
A director of Company X is designated.
Company X appears in a dataset that also contains sanctioned entities.
Company X has a commercial relationship with an entity that appears on a risk list.
Do not compress all of them into:
Company X is sanctioned.
The relation type is part of the finding.
PEP is not a synonym for criminal or sanctioned
A politically exposed person classification is a risk-management category used in financial-crime compliance contexts.
It is not an accusation of wrongdoing.
Likewise:
- relative or close associate;
- state-owned enterprise;
- sanctioned entity;
- sanctions-linked entity;
are distinct concepts.
If your source uses a semantic risk topic, preserve the exact topic.
Avoid moral language the data does not support.
Trace risk data back to the underlying source
Aggregators are valuable because they normalize many lists.
For important findings, inspect:
- source dataset;
- issuing authority;
- program;
- designation date;
- identifiers;
- current status;
- any removal or amendment history.
The strongest statement is tied to the authority.
For example:
OpenSanctions aggregates record X from authority Y.
Then verify authority Y when practical.
Step 7 — Use documents to answer relationship questions
Corporate research frequently turns into document research.
Useful documents can include:
- annual accounts;
- incorporation documents;
- shareholder filings;
- official notices;
- procurement records;
- court filings;
- regulatory actions;
- contracts;
- leaked documents.
Documents need provenance.
Record:
Document title
Source
Document date
Publication date
Pages
Entity identifiers
Relevant passage
Interpretation
A company name appearing in a document proves only that the document contains that name.
The significance depends on context.
Document mention is not participation
Suppose a leaked spreadsheet contains:
Example Holdings Ltd
Possible interpretations include:
- customer;
- supplier;
- owner;
- target;
- unrelated reference;
- copied address book;
- historical record.
The document itself must establish the relation.
Avoid:
Company X was involved.
when the evidence only supports:
Company X is mentioned in document Y.
This is one of the most important writing disciplines in corporate OSINT.
Step 8 — Connect the company's claimed web identity
A company's website can help corroborate legal identity.
Useful first-party signals include:
- legal footer;
- registered address;
- company number;
- privacy notice;
- terms;
- contact page;
- corporate group statement.
Compare those against registry data.
Strong consistency
Website legal name = registry legal name
Company number matches
Address matches
This supports the relationship between the website and legal entity.
Weak consistency
Brand name only
No registration number
Generic contact form
The website may still be legitimate, but the entity linkage is weaker.
Contradiction
Website claims Company A
Terms identify Company B
Registry address belongs to Company C
That deserves investigation.
Do not automatically assume fraud.
Corporate groups, trading names and outsourced services can create legitimate complexity.
Historical websites can explain corporate change
The Wayback Machine can help answer:
- Did the site previously identify another legal entity?
- Did the registered address change?
- Did branding change around a merger?
- Did privacy terms name a previous company?
- Did a domain previously belong to another business?
Historical web evidence is especially useful when legal records show:
- name change;
- acquisition;
- dissolution;
- restructuring.
Archive evidence remains incomplete.
A missing capture does not prove absence.
Step 9 — Preserve evidence before the research expands
Company research produces many volatile sources.
Some can change:
- registry interfaces;
- websites;
- sanctions entries;
- press releases;
- online filings.
For important findings, preserve:
- URL;
- source name;
- retrieval time;
- identifiers;
- relevant text or record;
- local copy or capture where appropriate and lawful.
Tools such as Hunchly can help package investigative browsing and evidence trails.
Public archives can preserve web states.
These are different preservation models.
Choose the one appropriate to the investigation.
Step 10 — Build a structured evidence ledger
A useful company research ledger might contain:
| Claim | Evidence | Source | Date | Confidence | Notes |
|---|---|---|---|---|---|
| Company legally exists | registry record | official registry | date | high | number X |
| Website linked to entity | legal footer + company number | company website | date | high | matches registry |
| Person A is director | filing | registry | date | high | current as of date |
| Person A owns company | unknown | — | — | unresolved | director ≠ owner |
| Sanctions candidate exists | OpenSanctions candidate | aggregated source | date | medium | identity review required |
The ledger prevents weak signals from silently becoming facts.
Use confidence deliberately
A simple confidence model can help.
High confidence
Use when:
- authoritative identifier matches;
- official record is clear;
- multiple independent sources agree;
- relation type is explicit.
Medium confidence
Use when:
- evidence is strong but incomplete;
- structured aggregator and first-party data agree;
- document context supports the relation but one element remains unresolved.
Low confidence
Use when:
- name-only match;
- single unverified secondary source;
- weak address similarity;
- inferred relationship;
- stale or ambiguous record.
Do not hide low confidence.
Low-confidence findings can still guide the next step.
A worked example
Suppose a website claims to belong to:
Northstar Logistics
You need to determine the underlying company and check whether a sanctions-related claim circulating online is accurate.
Phase 1 — Resolve the company
The website footer says:
Northstar Logistics Ltd
Company No. 12345678
OpenCorporates returns a company with:
- matching name;
- matching jurisdiction;
- matching company number.
You follow the provenance to the official registry.
The registry confirms:
- company number;
- active status;
- registered office.
Finding
The website's stated legal entity is consistent with the official company registered under number 12345678 as of the checked date.
This is a strong identity finding.
Phase 2 — Officers
The registry names:
Person A
as a current director.
Record:
Person A is listed as a director.
Do not write:
Person A owns the company.
unless ownership evidence supports that.
Phase 3 — Sanctions screening
A search for:
Northstar Logistics
returns a similarly named sanctions candidate in another jurisdiction.
The candidate has:
- different registration number;
- different address;
- different country.
Finding
A similarly named sanctioned entity exists, but the available identifiers do not support treating it as the same legal entity.
This negative resolution is valuable.
You prevented a false positive.
Phase 4 — Investigative documents
Aleph contains a document mentioning:
Northstar Logistics Ltd
with the correct company number.
The document places the entity in a supplier list.
Finding
The entity is identified as a supplier in document Y.
Do not transform:
supplier
into:
owner, conspirator or controlled company
without additional evidence.
Phase 5 — Historical website
A Wayback capture shows the same domain previously listing a different legal entity before a documented acquisition.
Now the timeline has context.
Calibrated conclusion
A defensible summary could be:
Northstar Logistics Ltd, registration number 12345678, is the legal entity currently identified by the website and registry. Person A is listed as a director. A sanctions search produced a similarly named entity, but the available identifiers do not support identity. Historical web evidence indicates that the domain previously identified a different legal entity before the current corporate structure.
Every sentence has a different evidence basis.
That is what good company research should look like.
Common mistakes
Mistake 1 — Starting with the company name instead of the entity
Names collide.
Use identifiers.
Mistake 2 — Treating OpenCorporates as the registry itself
It is a structured source that can provide provenance and links.
Verify important facts against the underlying register when practical.
Mistake 3 — Treating every Aleph match as identity
Cross-reference results need entity-resolution review.
Mistake 4 — Treating a sanctions candidate as a confirmed match
Screening is matching.
Matching produces candidates.
Mistake 5 — Calling a PEP "sanctioned"
These are different categories.
Mistake 6 — Confusing director, owner and controller
Record the exact relationship.
Mistake 7 — Treating document mention as participation
A mention requires context.
Mistake 8 — Ignoring dates
Corporate structures change.
Mistake 9 — Building an unlabeled network graph
Every edge needs a source and relation type.
Mistake 10 — Researching every connected entity
Expand only when a relationship is relevant to the original question.
A repeatable company-research workflow
Use this sequence.
1. Define the question
What exactly are you trying to establish?
2. Resolve the entity
Collect:
- legal name;
- jurisdiction;
- registry number;
- identifiers.
3. Verify official status
Check the strongest available authoritative record.
4. Record names and identifiers
Preserve:
- current name;
- previous names;
- company numbers;
- LEI where relevant.
5. Map explicit relationships
Directors, owners, parents, subsidiaries or other relationships only when supported.
6. Search document and investigative sources
Use entity keys, not names alone.
7. Screen sanctions and risk data
Only after identity is stable.
8. Review the underlying risk source
Determine:
- direct listing;
- relationship;
- topic;
- date;
- source authority.
9. Corroborate the company's web identity
Compare first-party claims to legal records.
10. Add historical context
Use archival and dated records when the question requires a timeline.
11. Preserve evidence
Keep provenance before interfaces or pages change.
12. Write facts, inferences and unresolved questions separately
Then stop when the research question has been answered.
Related OSINT.dev tools
OpenCorporates
Use for cross-jurisdiction company discovery and structured company records.
Treat provenance and official-source links as part of the result.
OCCRP Aleph
Use when the research moves into:
- large document collections;
- cross-referencing;
- entity relationships;
- investigative datasets.
Resolve entity identity before interpreting matches.
OpenSanctions
Use for sanctions, PEP and related risk-data screening.
A matching result should be treated as a candidate that requires identity validation and source review.
Wayback Machine
Use historical public pages to compare:
- legal footers;
- company names;
- group ownership statements;
- addresses;
- branding.
Hunchly
Use case-oriented capture when the investigation requires a reproducible browsing and evidence trail.
The core principle
Responsible company research is not a search for suspicious connections.
It is a process of identity resolution and evidence control.
Use this order:
question → legal identity → authoritative record → explicit relationships → documents → risk screening → corroboration → preservation → calibrated conclusion
The order protects the investigation from its most dangerous error:
finding a compelling story before establishing that the records refer to the same entity.
A name is a lead.
An identifier is stronger.
A relationship needs a source.
A risk match needs validation.
A document needs context.
And every conclusion should be traceable back to the evidence that supports it.
References
Primary and official product/data documentation used in this guide:
-
OpenCorporates — Company data and API
https://api.opencorporates.com/ -
OpenCorporates — Data provenance documentation
https://knowledge.opencorporates.com/ -
OCCRP Aleph — User Guide
https://docs.aleph.occrp.org/ -
OCCRP Aleph — Cross-referencing your data
https://docs.aleph.occrp.org/users/investigations/cross-referencing/ -
OpenSanctions — Understanding OpenSanctions Data
https://www.opensanctions.org/docs/data/ -
OpenSanctions — Matching API
https://www.opensanctions.org/docs/api/matching/ -
OpenSanctions — Using PEP Data
https://www.opensanctions.org/docs/pep/using/ -
GLEIF — Global LEI Index
https://www.gleif.org/lei-data/global-lei-index -
GLEIF — LEI Reference Data: Who Is Who
https://www.gleif.org/en/lei-data/access-and-use-lei-data/level-1-data-who-is-who
OSINT.dev · Published Apr 21, 2026 · Updated Aug 18, 2026. Canonical URL: https://osint.dev/articles/a-responsible-method-for-company-research-with-public-sources
Related articles.
Editorial pieces that share a tool context or type with this one.
How to Use Sanctions and Risk Lists Without Overreading Them
A disciplined method for sanctions, PEP and risk-list research: resolve the entity, classify the signal, verify the authority, distinguish ownership and control rules, preserve time, and avoid turning adjacency into a verdict.
Hunchly vs ArchiveBox: Evidence Packaging vs Archive Ownership
Compare Hunchly and ArchiveBox as two different preservation operating models: investigator-centered evidence capture, hashing and reporting versus self-hosted, multi-format archive ownership and recurring URL preservation.
OpenCorporates vs Aleph: Which One Fits Which Research Job?
Choose OpenCorporates when legal-entity identity and provenance are the main uncertainty; choose OCCRP Aleph when a known entity needs cross-dataset, document and relationship context.
A Practical Method for Domain and Infrastructure Recon
A passive-first, layer-by-layer workflow for domain and infrastructure reconnaissance using DNS, certificate transparency, HTTP behavior, technology signals, historical context and broader internet observations without turning discovery into attribution.