- AI can surface supply chain risk before it hits financial statements. Analyzing changes in 10-K and 10-Q language can reveal early warning signals across hundreds of suppliers.
- The opportunity is scale but not replacing analyst judgment. AI can continuously screen SEC filings and flag meaningful changes for subsequent deeper fundamental analysis.
- Case Study: Spirit AeroSystems demonstrates the potential. The system flagged deteriorating disclosure language 11 months before Boeing’s January 2024 door-plug incident.
A buy-side analyst covering dozens of stocks is also responsible for monitoring the financial health of hundreds of suppliers. Unless they are private companies, suppliers typically file quarterly and annual reports with the SEC. Taken together, the process of monitoring the financial health of suppliers comprises tens of thousands of pages a year and a significant time commitment. No analyst can cover their whole portfolio of companies with the same level of analytical rigor that they dedicate to their key coverage stocks.
As a result, portfolios face additional risk that arises from inattention to the financial standing of suppliers of lower priority holdings. Often the warning signs were there months before it manifested in their public income statements. This is a volume problem. Volume is the one constraint that even seasoned analysts cannot solve. An experienced analyst’s edge, with sector knowledge, management access, and judgement built over years, is something an AI model cannot yet meaningfully challenge. This edge, however, does not stretch to hundreds of supplier filings a quarter. Processing text at scale is one task where a computer genuinely outperforms an analyst due to speed and continuous processing.
Catching the Signal Before it Appears in the Numbers
Decades of accounting research confirm that warnings often arrive in the language before it is realized in the numbers. Firms facing challenges tend to write disclosures that are longer and vaguer, sometimes bogging down in dense complexity before the results turn negative. The signal is public, yet often uninterpreted. This is the gap we set out to close.
Our system pulls 10-K and 10-Q filings directly from the SEC’s EDGAR database across an initial universe of public companies in the US Industrials sector. From each company’s filing, we extract the Management Discussion & Analysis (MD&A) and Risk Factor sections, where disclosure language changes the most. Before scoring these items, we classify each sentence within them into predefined risk topics, including regulatory, geopolitical, and macroeconomic risks. Through this classification, our system understands what a company is discussing before addressing how that company is responding.
The Engine: NLP, Knowledge Graphs, and Targeted LLMs
At the core of our system are two language models, each fine-tuned to score all financial text on two dimensions. The first measure is vagueness—is the company being specific with its language or hedging? The second measures complexity—is this a genuine technical disclosure, or is bad news being buried in complexities? The two-pronged system is purposeful; a company’s challenges can be wrapped in vagueness or complexity, sometimes both. That is why the score of a single variable cannot tell you enough.
What matters most is deviation from a benchmark. We benchmark every company against its sector peers and against its own filing history, identifying declining trends and sector outliers. Evaluating from absolute scores can present bias in the results of our models. By measuring deviation from peer averages instead, we provide a safeguard which cancels out the potential of any bias.
These classified sentences populate a knowledge graph connecting each company to its industry peers, their filings, and its own filing history. This lets the system move beyond asking whether a risk factor section has changed at all, to a more precise question: has the company's disclosure on a specific risk topic shifted, relative to both its peers and its own prior filing? A lightweight model makes the first interpretive pass over the extracted sections, working alongside the vagueness and complexity scores. At roughly 97% lower cost per token than a frontier model, it is cheap enough to run across the defined universe. Only where that first pass identifies a genuine shift is the filing escalated to a frontier model for the deeper read.
What Lands on an Analyst’s Desk
The analyst receives a single memo for each flagged supplier: a concise analysis of what changed, the underlying scores and their trend against sector peers, as well as every claim traced back to the filing section that produced it.
We also developed a dashboard application to visualise the data for analysts. The dashboard brings together the generated memos, a trend explorer for investigating movements in a company’s or its sector’s vagueness and complexity scores, and a geographic heatmap which shows disclosure trends within a specific country.
In Practice: The Spirit AeroSystems Case Study
To test this signal, we used our system to investigate Boeing’s supply chain in the period before the January 2024 door-plug blowout. Spirit AeroSystems, one of Boeing’s largest fuselage suppliers, was operationally fragile long before the critical safety incident. To validate our test, we omitted certain data that would prejudice the results. We ran the system using the scoring models alone and excluded the LLM layer entirely because those models are trained on public text, which includes coverage of what eventually happened to Boeing and Spirit. Therefore, implementing this feature would introduce a look-ahead bias, where the models reflect what they know the outcome will be rather than what the language showed at the time.
Run in early 2023, our system escalated Spirit eleven months before the incident, and the multi-billion dollar fall in Boeing’s market value that followed. At any single point in time, Spirit's language scores were not extreme outliers relative to its peers, allowing it to hide in the noise. The signal was in the drift, with language moving steadily in the wrong direction relative to its sector peers over time. This highlights the importance of ongoing real-time analysis of a company's own filings.
Future Frontiers
The biggest takeaway from this exercise wasn’t about predicting a supply chain failure. It was the realisation that AI’s best use case in fundamental analysis isn’t replacing human judgement but directing analyst attention when volume and priority can drown out signals.
The models remain imperfect; interpreting genuine complexity is harder than vagueness and our complexity model training data was of lower quality than the vagueness model. Since submitting the system to the CFA AI Investment Challenge in April 2026, we have already improved the F1 scores of our models by implementing discriminative fine-tuning. Next steps are to explore developing these models further.
In our submission, we tested this system on US Industrials, but the method is sector agnostic. Wherever companies file structured disclosure, the same signal holds. The frontier we find most compelling is exploring earnings call transcripts, particularly the Q&A sections of these events, where management's language can be less prepared and more indicative. This represents an area where the strongest signals remain undetected.
If you liked this post, don’t forget to subscribe to the Enterprising Investor.
All posts are the opinion of the author. As such, they should not be construed as investment advice, nor do the opinions expressed necessarily reflect the views of CFA Institute or the author’s employer.
Image credit: ©Getty Images