Coverage

Mapping a fragmented industry from public data

Define the market for your acquisition thesis, assemble a target universe from public sources, and assess how much of that universe your research covers.

Updated 2026-09-17 · 2,877 words · 19 sources · jump to sources

A market map defines the universe behind a target list. It brings together company identities, business activities, ownership, and geographic coverage, then compares the identified companies with a reasoned estimate of the wider market. Public statistics, sector registers, and company research each contribute a different part of that assessment.

Why the map comes first

A market map helps answer three questions relevant to an investment committee.

  1. How many companies are there? Without a denominator you cannot say whether your twelve targets are the twelve best or the twelve you happened to find.
  2. What fraction have we looked at?“We have reviewed the market” is a claim about coverage, and coverage is a ratio. State the estimated universe and the basis for that estimate.
  3. How concentrated is the market? Assess the share held by major operators, the measure used, and the limits of the available data.

A maintained map also preserves research across changes in personnel. Recording what each operator does, who owns it, and how the information was established makes that knowledge available to the team.

NAICS, and where it lets you down

Nearly every public dataset keys on the North American Industry Classification System [1], so a map starts by choosing codes. This step is where most maps go wrong, and the failures are predictable.

FailureWhat happensMitigation
The colloquial industry is not a code'Managed IT services' spans several codes, none of which is it. Picking one drops most of the market.Start from the companies you already know and look up how each is actually classified, then take the union.
Self-classification driftFirms choose their own code on registration and often choose badly or never update it after pivoting.Never rely on a single code as a filter. Use codes to bound the search, then verify by reading what the company says it does.
The code is too broadYour code contains ten thousand establishments, most of which are the wrong kind of business entirely.Layer a second axis: size band, licence status, geography, or a capability the target must have.
The code is too narrowYou get a clean count that excludes the adjacent firms most likely to be acquirable.Map the neighbours deliberately and decide in writing which are in scope, rather than letting the code decide for you.
NAICS is a necessary coordinate system and a poor definition of a market.

The discipline that helps is to write down the capability test separately from the codes: a sentence saying what a company must actually do to be in your market. The codes then become a way of finding candidates, and the capability test is what decides whether each one is in. Confusing the two is how a map ends up containing firms that share a code and nothing else.

Establishing a denominator

Once you have codes, the federal statistical system will tell you roughly how many businesses should exist. Three sources, measuring three different things, and the differences between them are informative rather than annoying.

  • County Business Patterns counts establishments — physical locations — by code and county, with employment size classes [2]. Best for geographic structure.
  • Statistics of U.S. Businesses counts firms by enterprise size, aggregating all establishments under common ownership [3]. This is the one that answers “how many companies could we buy.”
  • The Economic Census, every five years, adds receipts and concentration ratios [4], which is the only public read on revenue distribution within an industry.

Then QCEW gives near-census employment and wages by county and industry from administrative records rather than a sample [6], and Business Employment Dynamics gives you the churn rate: how many establishments in this sector open and close in a normal year [7]. Churn is the base rate your thesis competes with, and it is routinely ignored.

All of these are APIs, not PDFs [5]. Pulling the cut you actually need — your codes, your size band, your states — takes an afternoon, and the number you get will differ from the number in any published summary, because no published summary used your cut.

If you want a sanity check on the economics rather than the counts, aggregate financial data by industry and entity type is published from tax returns [8]. It will not tell you about a company, but it will tell you quickly whether the margin you have penciled in for the sector is plausible.

Finding the near-census

Statistics give you counts. They do not give you names. For names, the highest-value move is to find a register that enumerates the industry, and in regulated sectors one usually exists.

The clearest example is telecommunications: providers of interstate telecommunications file an annual worksheet, which makes the filer list close to a complete enumeration of the industry [9], while licences, assignments and transfers of control are separately published [10]. A team with the filer list is working from the population. A team with a purchased list is working from a sample nobody has characterised.

The same pattern recurs across the economy. The question to ask is always the same:

What can an operator in this industry not legally do without somebody’s permission? Find who grants that permission. They publish a list.
  • Transport: operating authority and safety registration for anything moving freight across state lines.
  • Healthcare services: state facility licensure, accreditation bodies, provider enrolment files.
  • Financial services: state insurance department producer registers, adviser registration.
  • Industrial: environmental discharge and air permits, which name the operator and the site.
  • Construction and trades: state contractor licensing boards, usually with licence class and date of first issue.
  • Anyone selling to government: entity registration [11] and the award record [12].

Where no register exists, the fallbacks are trade association membership lists, trade show exhibitor directories, certification and partner directories published by equipment vendors, and industry award shortlists. Each is a biased sample — membership costs money, so the smallest operators are missing — but several biased samples with different biases, unioned, cover a great deal.

Building the numerator

With a denominator and one or more candidate registers, the map becomes a merge. In practice:

  1. Union every candidate sourcebefore filtering anything. Filtering early bakes each source’s bias into the result permanently.
  2. Resolve each candidate to a legal entity and a website. This is the expensive step and it is not optional; everything downstream depends on knowing that two rows are the same company.
  3. Apply the capability test by reading what each company says it does, not by trusting the code it was filed under.
  4. Record why each candidate was excluded. An exclusion without a reason is indistinguishable from an oversight, and the exclusions are where a map is challenged.
  5. Compare the surviving count to the statistical denominator and explain the gap. Every map has a gap. A map that claims none is not finished, it is unexamined.

Commercial databases belong in step one and nowhere else. They are good candidate generators and poor populations: coverage of small private firms varies systematically with size and country, which is documented carefully in the literature on building representative firm-level data [19]. Use them to add names, never to define the universe.

The deduplication problem

Merging sources creates duplicates, and duplicates are not a cosmetic problem. They inflate the universe, they split evidence across two records so that neither looks complete, and they cause the same owner to be contacted twice by two people, creating avoidable confusion.

Common identity issues include:

  • Name matching. Names can be shared across states and differ from legal names. Use names to identify possible matches, then confirm them with registry identifiers, addresses, ownership relationships, and website evidence.
  • Shared websites. Several legally distinct entities running one brand site. Whether that is one target or several is a judgement, and it has to be made explicitly and recorded, because the answer changes the count.
  • Acquired brands still trading. A site that looks like an independent company and belongs to a group. Check the current ownership before treating it as a separate target.
  • Holding company versus operating company. Both are real, both appear in registries, and only one of them is the business. An LEI record, where one exists, states the relationship explicitly [15]; otherwise it is a matter of reading filings [14] [13].

Identity resolution deserves to be a named, owned step with its own quality measure, not something that happens implicitly inside an import script. The measure we use is blunt: how many records had to be merged after the fact? A rising number means the resolution step is under-built.

Measuring concentration

Market concentration can be measured. The Herfindahl-Hirschman Index sums the squared market shares of all participants [16], and the competition agencies publish the thresholds at which they regard a market as concentrated [17].

You usually cannot compute a true HHI for a private-company sector, because shares require revenue you do not have. Two workable approximations:

  • Employment-based shares from your own map, using confirmed headcounts. Crude, and directionally informative if headcount correlates with revenue in the sector.
  • Published concentration ratios from the Economic Census, which reports the share of receipts held by the largest firms in an industry [4]. Lagged and coarse, and it is real data rather than an assumption.

Either way, state the measure and its basis rather than the adjective. “The four largest operators we identified account for roughly a fifth of mapped employment” is a claim somebody can check. “Highly fragmented” is a mood.

Adjacency, and where a market stops

The hardest judgement in mapping is not finding companies, it is deciding which ones are in. Markets do not have edges; classification systems pretend they do, and a map that adopts the pretence inherits an arbitrary boundary somebody else drew for statistical convenience.

Three tests, applied in order, do most of the work:

  1. The capability test.Does the company do the thing, with evidence? Not “mentions the thing” — a dedicated page, a price list, a case study, a certification. The distinction between a service a company sells and a service a company names is the single most useful line in a market map, and it is the one most often blurred by keyword matching.
  2. The substitution test. Would a customer of a company already on your map plausibly buy from this one instead? If yes, it is in the market even when the classification code says otherwise.
  3. The acquisition test. If your platform bought it, would the combination make sense to a customer? This is the one that matters for buy-and-build and it is the most permissive of the three, which is why it comes last.

Record the answers, including for the companies you exclude. An exclusion with a reason can be revisited when the thesis widens; an exclusion without one is indistinguishable from an oversight, and it is precisely the population an investment committee will ask about.

Evidence-weighted capability

A refinement worth the effort: rather than recording capability as a yes or no, record it with the strength of the evidence behind it. A company with a dedicated page, two case studies and a price list is a different proposition from one that lists the service in a footer. We distinguish three states — served (evidence at the level of a dedicated page, product page or case study), pursued (a certification or badge but no delivery evidence), and mentioned (the site names it and nothing more) — and only the first counts towards the map or the score.

If the capability definition changes, review the existing records against the revised standard. Track which companies still need that review so that the map does not silently combine assessments made under different definitions.

Keeping the map alive

A market map needs a refresh policy. The following intervals are illustrative starting points; adapt them to the mandate and update sooner when a relevant change is reported:

FieldRe-checkWhy
Existence and website reachableQuarterlyAn unavailable site is a reason to check the company’s status.
OwnershipEvery six to twelve monthsOwnership changes can affect eligibility and the appropriate contact.
Leadership and tenureEvery six monthsCheap from public profiles, and a change here is a timing signal.
HeadcountEvery six to twelve monthsNoisy quarter to quarter; a year-on-year move is meaningful.
Services and positioningAnnuallySlow-moving, and usually only matters when it changes materially.
Locations and capacityAnnually, or on any announcementPhysical facts change slowly and are usually announced.
Illustrative refresh schedule. Set intervals according to the mandate and available sources.

The important design point: a re-check should be triggered by a rule, not by somebody remembering. And a re-check that finds nothing new is a successful re-check and should be recorded as one, with its date, so the next person does not redo it.

A worked example, start to finish

To make the method concrete, here is the sequence for an invented mandate: a buyer looking for regional operators of facilities that host other companies’ equipment. This is an illustrative example of the method.

  1. Write the capability test first, in one sentence.“Operates at least one facility, under its own control, where third-party customers place equipment under contract.” Note what this excludes: brokers, resellers, and companies that merely resell somebody else’s capacity. Writing it before looking at data is what stops the data from defining the market for you.
  2. Take twenty companies you already believe are in the market and look up how each is classified. Check whether relevant companies span several codes and record the ones that apply [1].
  3. Pull the denominator for those codes, your size band and your states from firm counts rather than establishment counts [3], and separately pull establishments by county to understand the geographic shape [2]. Write both numbers down. They will differ by a factor that tells you how multi-site the sector is.
  4. Find the register. Ask what this operator legally cannot do without permission. Regulated communications infrastructure has filer lists and licence records [9] [10]; anything selling to government has registration and awards [11] [12]. Take everything.
  5. Union every candidate source without filtering. Registers, association directories, vendor partner listings, trade press, and commercial database results [19]. Record which source each candidate came from; you will want to know later which sources earned their keep.
  6. Resolve identity. One row per real operating company, with a confirmed website and a legal entity. Expect this to collapse the list noticeably and expect that collapse to be the most valuable step in the process [14] [15].
  7. Apply the capability test by reading, not by matching. Does the site show a facility, a customer, a contract? Record the evidence and record the exclusions with reasons.
  8. Compare to the denominator and write down the gap.“We have identified 340 against a plausible universe of 500 to 700; the gap is concentrated in single-site operators below the employee threshold in rural counties.” That sentence is what a market map is for, and almost no bought list can produce it.

Keep the capability definition and coverage estimate alongside the company records. They explain what the map includes, what it excludes, and where further research may be useful.

Knowing when a map is finished

It never is, but there is a defensible stopping point, and it is reached when you can answer these five questions without new work:

  1. What is the universe? A number, a definition, and the source of the denominator.
  2. What fraction have we identified by name? A ratio, with the gap explained.
  3. What fraction have we researched to the standard? A second, smaller ratio. These are different numbers and conflating them is how coverage gets overstated.
  4. Which excluded companies were excluded for which reason? By name, and reversible if the thesis changes.
  5. When was each record last confirmed? Per field, not per company.

These answers make the map usable for review and future updates. Keep the definitions, sources, coverage gaps, and refresh schedule with the company records.

Sources

References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.

  1. [1]
    North American Industry Classification System (NAICS)
    U.S. Census Bureau

    The classification every federal dataset keys on, and the first thing a market map has to get right.

  2. [2]
    County Business Patterns (CBP)
    U.S. Census Bureau

    Establishment counts, employment and payroll by NAICS code and county. The spine of any honest market map.

  3. [3]
    Statistics of U.S. Businesses (SUSB)
    U.S. Census Bureau

    Firm counts and employment by enterprise size, which is how you size the population of buyable companies.

  4. [4]
    Economic Census
    U.S. Census Bureau

    Every five years: receipts, establishments and concentration ratios by industry.

  5. [5]
    Census Data API: available datasets
    U.S. Census Bureau

    The machine-readable side of every dataset below. A free key, then plain HTTP and JSON; most people quoting census numbers have never called it.

  6. [6]
    Quarterly Census of Employment and Wages (QCEW)
    U.S. Bureau of Labor Statistics

    Near-census employment and wage counts by county and industry, monthly granularity, published quarterly.

  7. [7]
    Business Employment Dynamics
    U.S. Bureau of Labor Statistics

    Quarterly establishment births, deaths, expansions and contractions.

  8. [8]
    SOI Tax Stats
    U.S. Internal Revenue Service

    Aggregate financial data by industry and entity type, useful for sanity-checking a margin assumption.

  9. [9]
    Telecommunications Reporting Worksheet (Form 499)
    U.S. Federal Communications Commission

    Every interstate telecom provider files one, which makes the filer list a near-census of the industry.

  10. [10]
    Universal Licensing System (ULS)
    U.S. Federal Communications Commission

    Licence holders, assignments and transfers of control. A regulated industry publishes its own ownership changes.

  11. [11]
    SAM.gov entity registration
    U.S. General Services Administration

    Any entity that wants federal money registers here, exposing legal name, address and NAICS codes.

  12. [12]
    USAspending.gov
    U.S. Department of the Treasury

    Federal contract and grant awards by recipient. A revenue line you can read without asking.

  13. [13]
    EDGAR company and filing search
    U.S. Securities and Exchange Commission

    The public search interface. Free, rate-limited, and authoritative for anything a registrant had to disclose.

  14. [14]
    OpenCorporates
    OpenCorporates

    Aggregated company registry data across jurisdictions, with provenance back to the filing source.

  15. [15]
    Legal Entity Identifier (LEI) search
    Global Legal Entity Identifier Foundation

    Open, free entity identifiers with parent relationships. The nearest thing to a global primary key for companies.

  16. [16]
    Herfindahl-Hirschman Index
    U.S. Department of Justice

    The concentration measure the agencies use, and a discipline for claiming an industry is 'fragmented'.

  17. [17]
    2023 Merger Guidelines
    U.S. Department of Justice and Federal Trade Commission · 2023

    How the agencies say they analyse a transaction, including concentration thresholds and serial acquisition.

  18. [18]
    Premerger Notification Program (Hart-Scott-Rodino)
    U.S. Federal Trade Commission

    Filing thresholds are adjusted annually; below them a deal closes without notifying anyone.

  19. [19]
    How to Construct Nationally Representative Firm Level Data from the Orbis Global Database
    National Bureau of Economic Research · 2015

    A careful account of what commercial company databases are missing and how their coverage skews by size and country. Read before trusting any vendor's universe count.

Keep reading

This is how Docket works, not just what we think.

Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.