How ETFs Hide Your True Exposures
The name on an ETF is a marketing decision. Benchmark pressure and crude classification rules mean what you own can look very little like the label.
By Slava Tarasov, Co-founder
Buy a fund called "European Value," a second called "Global Dividend," and a third called "Artificial Intelligence & Robotics," and it is reasonable to believe you have built three different investments. You have diversified across style, across income, across theme. The names say so.
The names are marketing. What determines your actual exposure is a few hundred lines of holdings underneath each name, and those holdings are shaped by forces that have surprisingly little to do with the label on the fund. Two of those forces do most of the damage. The first is benchmark pressure: the professional reality that a fund manager who trails the index gets fired, which pushes even specialised funds toward whatever is currently driving the index. The second is classification methodology: the crude, single-bucket rules that decide whether a company counts as "retail" or "technology," "Europe" or "global," "value" or "growth" — rules that fail precisely on the companies large enough for the failure to matter.
This article walks through both forces, shows how they compound in a realistic three-fund portfolio, and ends with what you can actually do about it. None of it requires believing that anyone is acting in bad faith. The system produces misleading labels while everyone in it follows the rules.
A note on numbers and names: this is an opinion piece, not research. Figures are approximate, and the fund examples are stylised composites rather than named products — deliberately, because the mechanics are the point, and they survive any reasonable version of the numbers.
The label is a claim, not a measurement
Start with what an ETF name legally is. In most jurisdictions, a fund name has to be "not misleading" and, under rules like the SEC's names rule, a fund whose name suggests a focus must invest at least 80% of assets in line with that focus. That sounds strict until you notice who defines the terms. If the fund's prospectus defines "artificial intelligence company" broadly enough, Apple qualifies. If "value" is defined relative to a growth-heavy universe, a stock trading at thirty times earnings can be a value holding.
The definitions live in the index methodology or the prospectus, documents that almost nobody reads. The name lives on the fact sheet, the broker screen, and the portfolio summary — the places where decisions actually get made. So the practical situation is this: the most visible piece of information about a fund is the piece with the loosest connection to what it holds.
This matters because most investors assemble portfolios at the level of fund names. They hold five or ten funds and reason about diversification by reasoning about the labels: some tech, some dividends, some Europe, some bonds. If the labels were accurate summaries of the holdings, that would be a sensible shortcut. They are not, and the two sections that follow explain why.
Force one: benchmark pressure bends every fund toward the index
Every fund manager is measured against something. An active ETF has a stated benchmark, and the manager's performance — and career — is evaluated against it quarterly. Underperform for a few quarters and investors leave; underperform for a few years and the fund closes. This is not a hidden incentive. It is the explicit structure of the industry.
Underweighting a winner is the most dangerous position in professional investing. Suppose you run an active fund benchmarked against a broad index, and one stock — take Nvidia through 2023 and 2024 — comes to represent 6% or 7% of that index while tripling in price. If you hold none of it, and it keeps rising, you trail your benchmark by several percentage points from that single decision. Your investors do not see a principled stand. They see a manager who missed the biggest story in the market. The rational move, for anyone who wants to keep managing money, is to hold it anyway — at close to the index weight — whether or not it fits the strategy the fund was sold on.
This is how a dividend fund ends up holding stocks that barely pay dividends. The pattern shows up across categories. "Quality income" funds holding mega-cap growth names with sub-1% yields, because those names drive the reference index. European funds stretching definitions to include companies with US-listed lines. Value funds whose top ten look remarkably like the top ten of the growth index next door, because the manager could not afford another year of trailing while mega-cap growth ran. The industry has a name for the end state: closet indexing — a fund that charges active fees while hugging the benchmark so closely that the strategy in the name has almost no effect on returns. We will not do finger-pointing here — but anyone who watched value managers capitulate into mega-cap growth through 2020 and 2021 saw this mechanism play out in public, one quarterly holdings report at a time.
Tracking-error limits make it official. Many funds are explicitly constrained to stay within a band of the benchmark — a maximum tracking error of, say, 3% or 4% a year. Inside a constraint like that, when a handful of mega-caps make up a third of the index, holding those mega-caps is not a choice. It is arithmetic. The fund's stated philosophy operates only in the residual space the constraint leaves over, which in a concentrated market is not much space at all.
The consequence for you is direct: funds with different names converge on the same holdings. The style box says you diversified. The holdings say you bought the same seven companies four times.
Force two: classification puts trillion-dollar companies in one box
The second force is quieter and, at scale, worse. Nearly every exposure number you have ever seen — the sector pie chart, the country breakdown, the style box — is built on classification systems that assign each company to exactly one bucket.
The dominant one is GICS, the Global Industry Classification Standard, maintained by MSCI and S&P. It gives every listed company a single sector, industry group, industry, and sub-industry. One each. The assignment is based mainly on the company's principal business activity, with revenue as the primary measure and earnings as a secondary one.
For a regional bank or a copper miner, one bucket works fine. For the companies that dominate modern indices, it fails in a specific and instructive way.
Is Amazon a retailer or a cloud company? GICS has to pick one
Amazon is classified under GICS as Consumer Discretionary — sector-level "retail." Not Information Technology. The e-commerce business generates the large majority of Amazon's revenue, so under a revenue-first rule, retail it is.
But look at where the profit comes from. Amazon Web Services has in recent years produced on the order of 15–20% of Amazon's revenue while generating roughly 60% or more of its operating income. By economic engine — the thing that actually drives the share price on earnings day — Amazon is substantially a cloud-infrastructure company with a vast, thin-margin retail operation attached. Analysts value it that way. The market trades it that way: Amazon moves with Microsoft and Google on cloud and AI news, not with Target and Home Depot on consumer sentiment.
Now put that misfit into a portfolio. Amazon's index weight is in the low single digits of the S&P 500 — a top-five position. Every portfolio tool that relies on GICS reports that weight as "Consumer Discretionary." So:
An investor who wants to trim technology exposure looks at the sector chart, sees tech at some manageable number, and concludes there is no problem — while several percentage points of economically-tech Amazon sit hidden in the consumer bucket.
An investor who wants consumer exposure as a defensive tilt buys a consumer discretionary sector fund and finds that its largest holding is, economically, a cloud and AI bet.
Neither investor made a mistake. The measurement did. And at Amazon's scale — a company worth around two trillion dollars — a single-bucket answer to a two-business question moves whole percentage points of reported exposure into the wrong column. The error is not at the edge of the portfolio. It is at the centre.
Amazon is the cleanest example, but the same structural problem runs through the largest names in the market. Alphabet and Meta were classified as technology companies until 2018, when GICS moved them into the newly created Communication Services sector — overnight, index trackers' "tech" exposure dropped and their "communication" exposure jumped, with not a single share changing hands. Visa, Mastercard and PayPal were Information Technology until 2023, when GICS reclassified them into Financials. Tesla is Consumer Discretionary — an automaker — despite a market valuation that has always priced it substantially as a technology and energy company. In every case the companies did not change. The buckets did, or should have, and every sector-based exposure report inherited the distortion.
The lesson generalises: the bigger and more diversified a company becomes, the less meaning a single classification carries — and the bigger its index weight, the more that meaningless classification distorts your reported exposures.
Thematic ETFs: where both forces meet
Thematic ETFs deserve their own section because they combine both problems at maximum strength. A thematic fund's entire pitch is its label — clean energy, cybersecurity, robotics, AI — so the gap between label and holdings is the product, not a side effect.
Look inside a typical "AI" ETF and you will find Apple, Microsoft, Nvidia, Alphabet and Meta near the top. Are these AI companies? Partly, certainly. But you almost certainly already own them — they are the largest positions in any S&P 500 or global index fund, and in any Nasdaq fund, and in most active growth funds. The thematic fund is not giving you exposure to a theme. It is giving you a second, more expensive serving of the same mega-caps, with a sprinkling of small pure-play names underneath to justify the label.
The methodology documents reveal how this happens. Many thematic indices require only that a company derive some threshold of revenue — often 50%, sometimes as little as 25% — from the theme, or merely that it be "engaged in" the relevant activity. At a 25% threshold, most of the world's largest companies qualify for most technology-adjacent themes. The index provider needs investable capacity — a fund cannot put billions into micro-cap pure plays — so the methodology is written to admit the mega-caps, and the mega-caps then dominate the weight.
The result is a fund that charges 0.4–0.75% a year — several times the cost of a broad index fund — for a portfolio whose returns are driven mostly by companies you already hold at three to five basis points elsewhere. The theme in the name explains a small minority of the fund's behaviour. The rest is the same market beta you already owned, repackaged.
A worked example: three funds, one bet
Put the pieces together in a portfolio that looks sensibly diversified on paper. An investor holds, in equal parts:
- A global equity index ETF — the passive core.
- A "Quality Dividend" active ETF — for income and defensiveness.
- An "AI & Big Data" thematic ETF — a deliberate satellite bet on the theme.
Three funds, three distinct purposes, three different names. Now do the look-through.
The global index fund holds Microsoft, Apple, Nvidia, Amazon, Alphabet and Meta as its largest positions — together somewhere around 20–25% of the fund, reflecting their index weight. The dividend fund, benchmark-pressured as described above, holds Microsoft and Apple prominently: both pay dividends, both are "quality" by any screen, and no dividend manager benchmarked against a broad index could afford to skip them through the 2020s. The AI thematic fund holds Nvidia, Microsoft, Alphabet, Meta and Amazon near the top, because they clear the revenue threshold and provide capacity.
Sum it across the three sleeves and the same six companies plausibly account for 20% or more of the entire portfolio — a concentration the investor never chose and cannot see on any fact sheet, because each fund reports its holdings separately and no single document adds them up.
Now layer the classification error on top. Ask this portfolio's reporting tools "how much technology do I own?" and the answer will exclude Amazon (Consumer Discretionary) and, before the reclassifications, would have excluded Alphabet, Meta and Visa too. The reported tech number might say 28%. The economic answer — companies whose earnings rise and fall with cloud, devices, chips and digital advertising — might be north of 40%.
One correlated bet, wearing three names, mismeasured by the very chart that is supposed to reveal it. In a broad AI-led rally this portfolio looks brilliantly diversified, because everything goes up together and nobody questions rising numbers. The moment that trade cracks — a chip cycle turning, cloud spending pausing, a regulatory shock — all three funds fall together, and the investor discovers what they actually owned at the worst possible time to learn it.
Why this is worse now than it has ever been
Both forces have existed for decades. What changed is the market they operate in.
Index concentration is at generational highs. In 2015, the ten largest companies in the S&P 500 accounted for roughly 17–18% of the index. A decade later that figure has run well above 35%, and the top seven names alone have at times exceeded 30%. Global indices, which are supposed to dilute this, do not: US mega-caps are so large that they dominate world indices too, and the same handful of companies sit at the top of an MSCI World tracker, a Nasdaq fund and an S&P fund alike.
Concentration amplifies both distortions at once. Benchmark pressure gets stronger, because the penalty for skipping a mega-cap grows with its index weight — a manager could afford to skip a 1% position on principle, but not a 7% one. And classification error gets more expensive, because each misfiled company drags more portfolio weight into the wrong bucket. Amazon misclassified at half a percent of the index is a rounding error; Amazon misclassified at nearly 4% is a visible hole in every sector chart built on GICS.
Concentration also quietly rewrites what "the index" means. A cap-weighted index was sold to a generation of investors as the diversified, neutral choice. At today's weights, a broad US index fund is closer to a concentrated position in a handful of technology-adjacent franchises with several hundred small holdings attached. That may be exactly the bet you want — those franchises earned their weights — but it should be recognised as a bet. When the passive core itself is concentrated, and every active and thematic fund is pulled toward the same names by benchmark pressure, the diversification the labels promise has to be checked rather than assumed.
There is a plausible future in which this unwinds — concentration has mean-reverted before, and the investors who were paying attention in 2000 remember how violently the top of an index can deflate. Until then, the honest description of most fund-based portfolios is: more correlated than they look, more concentrated than they report, and labelled as if neither were true.
Geography has the same disease
Sector is not the only dimension where single-bucket classification misleads. Country classification is usually based on where a company is listed, incorporated, or headquartered — not where it earns money.
A "European equity" fund holds ASML, Novo Nordisk, LVMH, SAP and Nestlé. By listing, impeccably European. By revenue, these are global companies: ASML sells lithography systems predominantly into Asia; LVMH's growth story has for years been the Chinese consumer; the large-cap European indices as a whole derive well under half their revenue from Europe. An investor who buys a European fund to reduce dependence on the US economy and the dollar has done far less of that than the country pie chart suggests. The reverse also holds: an S&P 500 portfolio has, through the revenues of its constituents, very substantial exposure to European and Asian demand.
The pattern is the same as with Amazon: a one-bucket answer to a many-bucket question, accurate in form, misleading in substance, and most misleading for exactly the companies that dominate the weight.
What you can actually do about it
None of this argues against owning ETFs. Broad, cheap index funds remain the best vehicle most investors have ever had access to. The argument is against trusting labels, and the defence is mechanical rather than clever.
Read holdings, not names. Before buying any fund — especially anything active or thematic — open the full holdings list, not the top ten. If the top of an "AI" fund is the same five mega-caps as your index fund, you are not buying a theme; you are doubling a position. The methodology document, dry as it is, tells you the revenue threshold and answers the question "what is this fund forced to hold?"
Aggregate before you judge. Diversification is a property of the whole portfolio, not of any fund in it. Combine every fund's full holdings into a single list and look at your true top positions. The first time investors do this, the same discovery appears almost every time: the largest single-company exposure is two to four times what any individual fact sheet shows, spread across funds that were supposed to be different.
Treat sector and country charts as approximations. Any exposure number built on single-bucket classification inherits the Amazon problem. Where a company is large enough to matter, look at its segment economics — where the operating income actually comes from — and at revenue-based geographic exposure rather than listing-based. For the handful of trillion-dollar companies that dominate every index, this is not pedantry; it is the difference between a right and wrong answer about your biggest positions.
Watch for convergence over time. A fund that was differentiated when you bought it can drift toward the index as benchmark pressure does its work. Overlap between your funds is not a one-time check but something that grows silently in concentrated markets. The funds' turnover does the drifting; your job is to notice.
This is, ultimately, look-through analysis — the discipline institutional investors have applied for decades, decomposing every fund into its constituents and measuring exposure at the level of actual companies, actual segments, and actual revenue. There is no conceptual barrier to individuals doing the same; historically the barrier was data and tooling. That barrier is now falling, and the investors who take advantage of it will simply know what they own — which, in concentrated markets wearing diversified labels, is a genuine edge.
The ETF wrapper democratised access to markets. It did not democratise understanding of them. The label tells you what a fund is called. Only the holdings tell you what it is.
What a Balanced Investment Portfolio Actually Looks Like
Active vs Passive Investing: A Guide for Long-Term Investors
Are fund providers lying about what their ETFs hold?
No — and that is what makes the problem durable. Holdings are disclosed, methodologies are published, and names generally comply with regulation such as the 80% names rule. The distortion comes from structure, not deception: names are defined loosely enough to be marketing, benchmark pressure pushes different strategies toward the same mega-caps, and single-bucket classification misfiles the largest companies. Everything is technically accurate and the overall picture still misleads. That is why the burden of look-through falls on the investor: no single document anyone is required to publish shows you your aggregated, economically-classified exposure.
What is closet indexing and how do I spot it?
Closet indexing is when a fund charges active-management fees while holding a portfolio so close to its benchmark that the stated strategy has little effect on returns. The practical tells are a high overlap between the fund's top holdings and the index, a low tracking error (often disclosed in fund documents), and an "active share" — the percentage of the portfolio that differs from the benchmark — below roughly 60%. If a fund's active share is 30%, you are paying full active fees on the whole portfolio for active decisions applied to less than a third of it.
Why is Amazon classified as a retail company rather than technology?
GICS assigns each company to a single sector based primarily on where its revenue comes from, and the majority of Amazon's revenue still comes from e-commerce, so it sits in Consumer Discretionary. The distortion arises because Amazon's profits tell the opposite story: Amazon Web Services generates a clear majority of operating income from a minority of revenue. A revenue-first, single-bucket rule therefore classifies the company by its largest business rather than its economically dominant one — and because Amazon is a top-five index constituent, the misfiled weight is large enough to visibly distort the sector exposure of any portfolio that holds broad index funds.
Are thematic ETFs ever worth buying?
They can be, if you buy them for what they actually hold rather than what they are called. A thematic fund earns its place when its holdings are genuinely distinct from your core — pure-play companies you do not already own — and when you accept the concentration and fee that come with that. The test is mechanical: take the fund's full holdings, remove every position that already appears in your index funds, and ask whether you would pay the fund's fee for the remainder. Often the honest answer is no. Occasionally — usually with stricter revenue-purity methodologies and smaller constituents — it is yes.
Does this problem apply to plain index funds too, or only to active and thematic ETFs?
It applies to index funds as well, through a different door. A broad index fund holds exactly what it promises — the index — so there is no label-versus-holdings gap. But the exposure reporting built on top of it inherits every classification error: the sector chart of an S&P 500 fund files Amazon under consumer, filed Alphabet and Meta under technology until 2018 and communication services after, and reports country exposure by listing rather than revenue. And because cap-weighted indices are currently so concentrated, the index fund itself contributes the largest share of mega-cap weight to the overlap problem. The index fund is honest about what it holds; it is the summary statistics around it, and its interaction with your other funds, that mislead.
How much fund overlap is too much?
There is no universal threshold, but the question to answer is about companies, not funds: what percentage of your total portfolio sits in your single largest company across all funds, and in your largest five or ten? Many investors who run this exercise find 6–10% of everything in one company and 25% or more in the same handful of mega-caps, accumulated across funds with different names. Whether that is too much depends on your risk tolerance — but it should be a number you chose, not one you discover. If two funds share most of their weight, you are paying two fees for one exposure, and the more expensive fund needs a separate justification.