Data

Sources of acquisition target data and their limits

A field guide to registries, filings, company websites, and commercial databases: what each source can establish, where coverage is limited, and how to handle conflicting information.

Updated 2026-09-17 · 2,921 words · 24 sources · jump to sources

Acquisition research draws on sources with different strengths and limits. A registry can establish legal identity; a company website can describe its services; a commercial database can help identify candidates. This guide explains how to assess each source for the question at hand and retain enough context to review the resulting finding.

The premise: authority is per-field, not per-source

Assess a source against the specific field being researched. Its authority, date, and scope determine what it can support. Record these alongside the answer so a reviewer can assess conflicting findings.

Auditing offers a useful principle for assessing evidence: evidence obtained directly from an independent external source is more reliable than evidence produced by the entity under examination, and documentary evidence beats oral representation. Apply source preferences to individual research questions, taking account of relevance and date.

Company registries

Every jurisdiction maintains a register of legal entities. In the United States this is done at state level, which is the root of most of the difficulty.

SourceAuthoritative forSilent or weak onCharacteristic failure
State registries, e.g. Delaware [3]Existence, exact legal name, entity type, formation date, standing, registered agentOwnership, financials, trading names, what the business doesMost US states publish no shareholder information at all. Confirming an entity exists tells you nothing about who controls it.
OpenCorporates [4]Cross-jurisdiction entity lookup with provenance back to the source registerAnything the underlying register does not publishCoverage and freshness vary by jurisdiction because the upstream registers do. Treat it as an index, not as a source.
GLEIF / LEI records [5]A globally unique identifier, legal name, address, and direct and ultimate parent where reportedCompanies with no reason to have an LEI, which is most private firmsCoverage skews to entities with financial-market obligations. Superb when present, absent more often than not.
Companies House (UK) [6]Filed accounts, officers, and persons with significant controlCoverage and detail depend on the entity and filing requirementsDo not assume the same disclosures are available in other jurisdictions.
Registry coverage and disclosures vary by jurisdiction and entity type.

The practical consequence is that identity resolution is the first and hardest step, not a formality. A single operating business can present as a trading name on a website, a differently named holding company on a contract, three registered entities in three states and an acquired brand that still answers the phone. Any list that has not resolved these will double-count, and any score computed across a double-counted list is measuring your data pipeline.

Regulatory filings

Filings are the strongest public evidence available about private companies, because somebody was legally obliged to be accurate.

EDGAR, including the part about private companies

The obvious use of EDGAR is public companies. The valuable use, for origination, is that public companies are obliged to describe their private counterparties, acquisitions, customers and competitors, and all of that is full-text searchable [1]. Searching a private company’s name across filings frequently returns a merger background section describing a process it participated in, a customer concentration disclosure naming it, or a subsidiary list confirming who owns it.

Form D is another source: an exempt offering leaves a public notice with the issuer’s name and address and its executive officers and directors [2]. A private company that raised money is visible there even if it never appears in any press.

Intellectual property

Trademark records are the most underused ownership source in American research. A registration names an owner, an address and an entity type, and a trading name whose operator is otherwise opaque is frequently resolved in one lookup [7]. Patent assignment records capture changes of corporate control, often earlier than any announcement, because the paperwork has to be recorded whether or not anyone issues a release [8].

Procurement

Any entity that wants federal money registers, and that registration exposes a legal name, an address and self-declared NAICS codes [11]. Awards are then published with amounts and dates [12]. For a company with meaningful public-sector revenue, this is a partial income statement that you can read without asking anybody.

Sector registers

If your target industry is regulated or licensed, there is very often a register that amounts to a census of the industry, and it is the single highest-leverage thing you can find.

Telecommunications is the clearest example. Licence holders, and every assignment or transfer of control, are published [9]. Separately, providers of interstate telecommunications file an annual worksheet, which makes the filer list close to a complete enumeration of the industry [10]. An origination team that has the filer list has the universe; a team working from a bought list has a sample of unknown shape.

The same pattern recurs: state licensing boards for trades and professional services, department of transportation operating authority for logistics, state insurance department registers for brokers, environmental permits for industrial operations, health department registers for clinical services. Assess whether any applicable registers cover your intended market and how current their entries are.

Statistical agencies

Statistical sources cannot tell you about a company. They help assess the size of the market and the coverage of your target list.

  • NAICS is the classification everything else keys on [13]. Learn its boundaries before you build on it, because the code that describes an industry colloquially is often not the code its firms are classified under.
  • County Business Patterns gives establishment counts and employment by code and county [14], and SUSB gives firm counts by enterprise size [15]. Together they answer: how many companies of roughly this size, doing roughly this, should exist in this geography?
  • QCEW is near-census employment and wage data by county and industry [16], which is a sanity check on any claim about how large the sector is locally.
  • IRS SOI publishes aggregate financials by industry and entity type [17], which is how you find out that the margin you have assumed for a sector is a full standard deviation from what the sector reports.

The discipline these enable is coverage estimation: if CBP says there are roughly two thousand establishments in your codes and geographies and your list has three hundred companies, you know something specific about how much of the market you cannot see. Account for the distinction between establishments and firms before comparing those counts.

Self-reported sources

Professional networks, company profiles and directory entries are self-reported, which makes them excellent for some fields and dangerous for others.

FieldReliabilityWhy
Existence of a person in a roleGoodPeople are motivated to describe their own job accurately and are corrected by colleagues.
Tenure and job historyGoodSame. Dates are occasionally rounded, rarely invented.
Employee count on a company pageConditionalIt counts people who associated themselves with that page. Systematically understates blue-collar and older workforces, and is sensitive to which page is the right one.
Company description and industryWeakWritten for recruiting and search, not accuracy. Often years stale.
Revenue, ownership, founding dateVery weakUsually inferred by the platform or filled in once and never revisited.
Self-reported profile data: strong on people, weak on firmographics.

The employee-count row deserves emphasis because it is the field most often used as a hard filter. A headcount is only as good as the page it came from, and the failure mode is not noise but misattribution: a count read off the wrong company’s page, or off a parent’s page, is confidently wrong rather than approximately right. Any pipeline using headcount as a gate needs an explicit step that confirms the page belongs to the company, and needs to treat a count from an unconfirmed page as no answer at all.

Commercial aggregators

Commercial databases can help build an initial list and cross-check findings. Evaluate their coverage, field definitions, and update practices against your mandate.

  1. Coverage is not uniform and the gaps are not random. Careful work on constructing nationally representative firm-level data from a large commercial database found that coverage varies systematically with firm size and country, and that using the raw database as a population produces biased results [18]. That paper is about one product, and the lesson generalises: the smaller and more private your targets, the worse the coverage, which is exactly the segment origination cares about.
  2. Derived fields are models, presented as data. Estimated revenue, estimated employees and industry classification are frequently inferred. Check whether estimates are labelled and whether the provider explains its method.
  3. Everyone has the same screen. Anything you can filter, a competitor can filter identically. The output of a reproducible screen is not proprietary by construction.
  4. Staleness is invisible. A record shows you a value, not when it was last confirmed. A company acquired eighteen months ago can sit in a database looking independent.

An aggregator can serve as a candidate generator and a cross-check. For findings that materially affect a decision, seek primary support and record where independent confirmation is unavailable.

The company’s own website

For a private company the website is, surprisingly often, the best available source — with one very specific caveat.

It is authoritative for what the company says about itself: services offered, locations, named customers, certifications claimed, leadership, and the ownership statements that appear on about pages with remarkable frequency. It is a primary source, in the sense that the company published it. Many sites also publish structured markup describing the organisation, which can help identify relevant fields [19].

The caveat: a website is marketing. It over-claims capability, it lists partnerships as though they were products, it keeps pages for services the company quietly stopped selling, and it does not update when the company is acquired. So a website is strong evidence of what the company wants to be known forand weak evidence of scale, ownership currency and revenue mix. The right response is not to discount websites; it is to be precise about which claim the page supports. “The company has a dedicated page describing this service, with two case studies” is a defensible finding. “The company offers this service” is an interpretation. Keep the source statement distinct from the interpretation.

Collecting public information at scale sits in a legal area worth understanding rather than guessing at. This is not legal advice and your counsel’s view governs.

  • The Computer Fraud and Abuse Act turns on authorisation, not on technique.The Supreme Court narrowed “exceeds authorized access” to a gates-up-or-down question rather than a question about purpose [22]. The Ninth Circuit’s treatment of scraping publicly available profile data addressed whether data open to the general public is “without authorization” at all [21].
  • Contract is a separate question from statute.A site’s terms of use may prohibit automated collection even where no criminal statute is implicated. Those are different exposures with different remedies, and conflating them produces both false comfort and false panic.
  • robots.txt is now a standard, not folklore. The exclusion protocol was standardised in 2022 [20]. Honouring it is cheap, it is the norm, and ignoring it is the kind of detail that turns up in a diligence questionnaire later.
  • Personal data rules apply to business contacts. A named individual at a company is a natural person, and European and UK regimes treat their business contact details as personal data. That is covered in detail on our page about outreach.

When two sources disagree

Disagreement is not an edge case, it is the normal state of a research file with more than one source in it. What separates a usable file from an argument is having decided in advance how disagreements resolve.

The four kinds, and what each one actually means:

KindExampleResolution
Different vintagesA 2021 press release says sponsor-backed; a 2026 about page says independently owned.Not a conflict. Both are true of their dates. The projection rule is 'most recent evidence of adequate rank wins', and both claims stay on file.
Different entitiesOne source describes the holding company, another the operating subsidiary.Confirm which entity each source describes before comparing the values.
Different definitionsHeadcount of 40 from one source, 130 from another.Usually employees versus contractors, or one site versus all sites. Define the field precisely, then re-ask both sources against the definition.
Genuine contradictionTwo sources of comparable rank, same date, same entity, incompatible values.Record both, flag the conflict for review, and preserve the reason for any resolution.
Illustrative conflicts. Check entity, date, and definition before selecting an answer.

The design consequence is that a research store should never overwrite. If a newer value replaces an older one in place, the first row of that table becomes indistinguishable from the fourth, and you have destroyed the evidence needed to tell them apart. Keep every observation with its date and source, and compute the current view as a projection with an explicit rule. The projection can then be re-run when the rule improves, which is impossible once you have overwritten.

Researching business contacts

Everything above concerns facts about companies. Facts about people at those companies behave differently enough to need their own treatment, and conflating the two is how research files go stale in the most damaging way.

  • Roles and company facts have different update needs.Confirm a person’s current role before outreach and set refresh intervals appropriate to each type of information.
  • An email address has a verification state, not a truth value. There are three meaningfully different states: confirmed by a provider that tested it, inferred from a domain pattern that matches other addresses at the company, and guessed. Displaying all three identically is how a sending domain gets burned, because guessed addresses bounce and bounces destroy reputation.
  • Preserve reviewed corrections. Manual corrections should be able to override automated data permanently, and the override should be recorded as an override rather than silently merged, so that the next automated refresh does not undo it.
  • Person data carries legal weight that company data does not. A named individual is a natural person with rights of access and deletion under several regimes, which means person records need to be separable from the rest of the file rather than scattered through free text.

The practical architecture that follows: keep people in their own store, keyed to the company; record the verification state of every address explicitly; let human corrections layer over machine data with the newest non-null value per field winning; and never let a contact record silently become the evidence for a claim about the company. An owner’s statement on a call should be identified as such, with the date and context, so it can be distinguished from documentary research.

Keeping source material available

The final property of every source above is that it changes without telling you. Pew’s measurement of how much web content simply disappears over a few years is sobering for anyone whose research file is a column of URLs [24]: a live link may no longer let a reviewer inspect the material used.

Three approaches provide different levels of access to the original material:

  1. Store the link. Cheapest, and it quietly rots. By the time somebody checks, the page is a 404 and the claim is unverifiable.
  2. Rely on a public archive. Better [23], and dependent on whether the page was captured, whether the capture is complete and whether the archive is reachable when you need it.
  3. Retain the page yourself, with the date. Store the content you read, not the address you read it at, and cite the retained copy. This is what makes a claim checkable in two years, and it is the only one of the three that is fully in your control.
A source you did not keep is a source you cannot show. The link is the pointer; the evidence is the page.

A working practice

Pulled together, the field guide becomes six rules.

  1. Resolve identity before spending on research. Confirm the entity before using findings to assess fit.
  2. Rank sources per field and write the rank down. Filing beats company statement beats self-reported profile beats aggregator estimate, for most fields, and where your order differs, say so explicitly.
  3. Find your sector’s register. Check whether it can help define the target universe and identify gaps in your list.
  4. Estimate coverage against statistics.Know what fraction of the plausible universe you have, and tell your committee. “We have looked at all of it” is almost never true.
  5. Keep the page, not the link. With the date, and the sentence.
  6. Record not-founds as findings.“Revenue: not disclosed anywhere, five sources checked, as at this date” is a real answer that stops five people re-doing the same search. A blank cell teaches nobody anything.

Sources

References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.

  1. [1]
    EDGAR company and filing search
    U.S. Securities and Exchange Commission

    The public search interface. Free, rate-limited, and authoritative for anything a registrant had to disclose.

  2. [2]
    Form D — notice of exempt offering of securities
    U.S. Securities and Exchange Commission

    Private raises leave a public trace here, including the issuer's address and the size of the offering.

  3. [3]
    Division of Corporations entity search
    State of Delaware

    Entity status and formation date for the jurisdiction most US holding companies use.

  4. [4]
    OpenCorporates
    OpenCorporates

    Aggregated company registry data across jurisdictions, with provenance back to the filing source.

  5. [5]
    Legal Entity Identifier (LEI) search
    Global Legal Entity Identifier Foundation

    Open, free entity identifiers with parent relationships. The nearest thing to a global primary key for companies.

  6. [6]
    Companies House
    UK Government

    Free filings, accounts and persons with significant control. Ownership research is trivially easier in the UK than the US.

  7. [7]
    Trademark Status and Document Retrieval (TSDR)
    U.S. Patent and Trademark Office

    Ownership, addresses and specimens of use. A trading name's real owner is frequently here and nowhere else.

  8. [8]
    Patent Assignment Search API
    U.S. Patent and Trademark Office

    Assignments record changes of corporate control, often before any press release. The public search host has moved more than once; the API catalogue entry is the stable door.

  9. [9]
    Universal Licensing System (ULS)
    U.S. Federal Communications Commission

    Licence holders, assignments and transfers of control. A regulated industry publishes its own ownership changes.

  10. [10]
    Telecommunications Reporting Worksheet (Form 499)
    U.S. Federal Communications Commission

    Every interstate telecom provider files one, which makes the filer list a near-census of the industry.

  11. [11]
    SAM.gov entity registration
    U.S. General Services Administration

    Any entity that wants federal money registers here, exposing legal name, address and NAICS codes.

  12. [12]
    USAspending.gov
    U.S. Department of the Treasury

    Federal contract and grant awards by recipient. A revenue line you can read without asking.

  13. [13]
    North American Industry Classification System (NAICS)
    U.S. Census Bureau

    The classification every federal dataset keys on, and the first thing a market map has to get right.

  14. [14]
    County Business Patterns (CBP)
    U.S. Census Bureau

    Establishment counts, employment and payroll by NAICS code and county. The spine of any honest market map.

  15. [15]
    Statistics of U.S. Businesses (SUSB)
    U.S. Census Bureau

    Firm counts and employment by enterprise size, which is how you size the population of buyable companies.

  16. [16]
    Quarterly Census of Employment and Wages (QCEW)
    U.S. Bureau of Labor Statistics

    Near-census employment and wage counts by county and industry, monthly granularity, published quarterly.

  17. [17]
    SOI Tax Stats
    U.S. Internal Revenue Service

    Aggregate financial data by industry and entity type, useful for sanity-checking a margin assumption.

  18. [18]
    How to Construct Nationally Representative Firm Level Data from the Orbis Global Database
    National Bureau of Economic Research · 2015

    A careful account of what commercial company databases are missing and how their coverage skews by size and country. Read before trusting any vendor's universe count.

  19. [19]
    Schema.org vocabulary
    Schema.org

    The structured markup many company websites already publish about themselves, and the cheapest correct read of a homepage.

  20. [20]
    RFC 9309 — Robots Exclusion Protocol
    IETF

    The standardised version of robots.txt, published in 2022 after twenty-eight years as a convention.

  21. [21]
    hiQ Labs, Inc. v. LinkedIn Corp.
    U.S. Court of Appeals for the Ninth Circuit · 2022

    On whether scraping a public profile violates the Computer Fraud and Abuse Act. The answer turns on authorisation, not on whether you used a browser.

  22. [22]
    Van Buren v. United States
    Supreme Court of the United States · 2021

    Narrowed 'exceeds authorized access' under the CFAA to a gates-up-or-down question.

  23. [23]
    Wayback Machine
    Internet Archive

    The public record of what a page said on a date, when you did not retain it yourself.

  24. [24]
    When Online Content Disappears
    Pew Research Center · 2024

    Measures how much of the web cited a few years ago is already gone. The argument for retaining the page, not the link.

Keep reading

This is how Docket works, not just what we think.

Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.