Internalization determines where retail orders are filled—but not whether they receive good execution.
Expanded regulatory disclosures will make broker-level execution quality easier to compare.
Managers should test price improvement, spreads, fill rates, and speed across different conditions.
When a retail investor presses “buy,” the interface flashes green and a fill confirmation appears almost instantly. The experience suggests direct access to a public exchange. But before a marketable retail order interacts with displayed liquidity, it is processed, classified, and often absorbed by wholesale operators outside the visible market.
Most retail orders do not travel directly to the New York Stock Exchange or Nasdaq. Instead, they are routed first to wholesalers, who apply sophisticated algorithms to assess whether the order could force the market to move against them. That classification affects whether they internalize the order, how they price it, and whether they route it elsewhere.
Internalization occurs when a broker-dealer or wholesale market maker fills a customer order within its own trading operation rather than routing it to a lit exchange or another trading venue.
It is a routing mechanism — not a verdict on execution quality.
For institutional investors, the question isn’t whether an order was internalized; it’s whether execution held up against comparable alternatives once size, liquidity, volatility, and order type are accounted for. At the portfolio scale, even small, recurring differences in price improvement or effective spreads can translate into meaningful aggregate costs.
Why This Matters Now
Retail order flow is both large and growing. Vanda Research estimated that US retail investors generated $5.4 trillion in equity and ETF trading activity in 2025, nearly 47% more than in 2024. The US Securities & Exchange Commission Rule 605 provides greater visibility into broker-level execution, but managers must look beyond headline averages to determine who benefits from internalization, under what conditions, and at what cost.
A study by Robert Battalio and Robert Jennings provides a useful benchmark for why this matters. Using proprietary data on marketable retail orders routed to wholesalers, they found that external liquidity was used to fill 28.6% of the shares in their sample. They also found that the value of price and size improvement was 6.5 times greater than the value of price improvement reported under Rule 605. The finding does not establish that internalization is inherently worse; it shows that the standard price-improvement statistic can capture only part of the execution outcome.
What to Measure
Two measures provide the core of the analysis, and they answer different questions. Effective spread compares your execution price to the midpoint quote at the moment the order arrives. It tells you how far the fill sat from the prevailing midpoint at that instant. Realized spread compares the execution with the subsequent midpoint at a defined post-trade horizon. It adds a post-trade view of how the fill compares once the market has had time to respond, which is a different thing from checking whether the initial price improvement was “real.”
Neither number tells you much by itself. Both need company: execution speed, fill rate, order type, the time of day the order went in, and the prevailing spread and volatility when it arrived. Leave any of that out and you are measuring only the easiest orders, not execution quality.
Compare Like with Like
This is the part that turns a checklist into an actual finding. Comparing a 3,000-share order against a 100-share order and concluding that internalization failed because the bigger order cost more tells you nothing useful. Size alone is expensive to execute.
The comparison that matters is a 3,000-share order against other 3,000-share orders of the same type, in similar names, under similar liquidity and volatility, split by execution or routing path where that path can actually be identified. These comparisons are most informative when they are matched within the same symbol, order type, liquidity, and volatility bucket, and time-of-day window.
Once you do that, a lot of the variation caused by order difficulty can be controlled for, and the routing path becomes a more plausible contributor to whatever gap remains.
An illustrative case makes this concrete. Let’s say a manager’s 100-share orders in liquid large-caps come in around 1.5 basis points of effective spread, while 2,000-share orders in the same names run closer to 7 basis points. That gap alone proves nothing — larger orders are just harder to fill.
The real question is what comparable 2,000-share orders of the same type look like when routed externally under similar conditions. If those land around 6.8 basis points, internalization isn’t obviously the culprit; difficulty explains most of it. If this 2,000-share order lands around 3.5 basis points instead, the routing model for that size and liquidity bucket deserves a closer look.
The Pattern That Should Concern You
One bad fill does not tell you much. What should get your attention is a gap that persists after you’ve done the matching above — execution that stays worse than comparable alternatives once you’ve accounted for size, liquidity, volatility, and order type, and that keeps showing up rather than disappearing when market conditions normalize.
Internalization on its own is not the warning sign; it’s the gap that survives after accounting for how hard the order was to fill that is.
Where the Data Take You, and Where They Stop
Public disclosures can point you toward something worth investigating, but they rarely let you pin a single order to a single routing path. Rule 605 reports give execution-quality statistics at the market-center level. Rule 606 reports show how a broker’s routing choices relate to payment-for-order-flow and internalization arrangements. Together they can surface a pattern worth chasing. Confirming that pattern at the order level usually takes transaction-level data from your own broker’s execution reporting, since that is what identifies the venue or routing path for each fill. Think of the public record as where you start looking, not where you stop.
Anyone citing these reports today should note the timing. Rule 605 was amended in March 2024, and the compliance date was August 1, 2026. Broker-dealers newly brought into the expanded reporting framework are only just beginning to report under it, with more realized-spread horizons and broader order-size categories than before.
The first monthly reports covering August 2026 activity are due by the end of September. A separate piece of this — price-improvement statistics measured against the best available displayed price, which folds in the best-priced odd-lot information — has its own compliance date of November 1, 2026. Anything cited before those dates still reflects the older reporting structure, and any comparison spanning both regimes should say so plainly rather than treat the data as one continuous series.
Best Execution Is Not an NBBO Test
Beating the National Best Bid and Offer (NBBO) on a fill is not, by itself, proof that best execution happened. FINRA’s framework under Rule 5310 asks firms to use reasonable diligence to find the best market and get as favorable a price as the conditions allow, weighing market conditions, order size and type, execution price, price improvement, likelihood and speed of execution, transaction costs, and routing arrangements.
It is not a single pass/fail benchmark. FINRA guidance is also clear that orders a firm determines to execute internally remain subject to best-execution obligations and are subject to an order-by-order analysis of execution quality. Internalizing an order does not put it under a lower bar.
The test is a comparison, not a checkbox.
What to Ask Your Broker
Six questions get you further than any standard disclosure ever will.
What share of my orders is filled internally versus routed out, broken down by order type and size?
How does the execution quality of my orders stack up against reasonably available alternatives for comparable orders?
What does my realized spread look like at defined post-trade intervals?
How does execution quality shift for larger orders and during volatile stretches, once you compare against similar orders rather than the average?
What benchmark or comparison set do you use when judging execution quality on my flow?
And how do routing arrangements, including payment for order flow (PFOF), influence the path my orders take?
That last question is not meant to assume the arrangements are hurting you — it is meant to find out whether and how they shape the routing decision.
If you liked this post, don’t forget to subscribe to the Enterprising Investor.
All posts are the opinion of the author. As such, they should not be construed as investment advice, nor do the opinions expressed necessarily reflect the views of CFA Institute or the author’s employer.
Image credit: ©Getty Images