This report explores how human and large language model (LLM) biases enter investment workflows and how investment practitioners can detect, measure, and mitigate this phenomenon through prompt engineering, guardrails, comparative testing, and human oversight.
At a Glance
The report does the following:
- Identifies the sources of bias by distinguishing human bias, implicit LLM bias, and the observable choices a model makes within an investment workflow
- Demonstrates framing bias in action by presenting how various LLMs evaluated identical financial information differently when researchers changed only the information’s positive or negative presentation
- Compares practical mitigation methods, showing that balanced prompts and human intervention outperformed bias-awareness instructions and how a separate detection model could provide a scalable guardrail
- Shows how artificial intelligence (AI) governance can be strengthened by outlining how comparative testing, structured prompt templates, checkpoints, internal benchmarks, and human oversight can make bias more detectable and manageable
Analysis and insight from CFA Institute Research and Policy Center, delivered to your inbox.
What "Managing LLM Bias in Investing" Is About?
LLMs are becoming part of investment research, portfolio analysis, risk management, and client service. Their speed and scale can improve productivity, but biased inputs, model behavior, and workflow decisions can also distort recommendations, amplify errors, and create financial, regulatory, ethical, and reputational risks.
“Managing LLM Bias in Investing: From Detection to Mitigation” explores how bias can influence AI-assisted investment decisions. It examines common human biases, such as availability, anchoring, framing, and positional and self-preference bias, and explains how these can interact with AI prompts, selected information, system instructions, model design, and AI systems that make decisions or take actions during a workflow (agentic AI workflows) to reinforce biased outcomes.
The report combines behavioral finance with original experimental research to help firms build more transparent and reliable AI-enabled investment processes. It distinguishes implicit LLM bias, which arises from pre-training data, model architecture, and training procedures, from explicit LLM bias, which appears in observable choices such as data selection, source use, and analytical steps.
This distinction shifts attention from whether a model is simply “biased” to how a complete investment workflow produces its result. That broader view helps firms locate the source of a problem, select an appropriate control, and assign responsibility for reviewing the final decision.
What Did the Case Study Find?
The authors tested whether positive or negative framing changed how LLMs evaluated identical financial information, sending 900 prompts to each of nine OpenAI models across 10 investment scenarios, including drawdown performance, client retention, credit defaults, concentration risk, stress testing, and portfolio performance. Framing most strongly and consistently affected evaluations of drawdown performance and client retention. Simply defining framing bias in the prompt did little to reduce it. Manually presenting both positive and negative perspectives — the “human-in-the-loop” method — produced the largest reduction. A separate bias-detection LLM also identified and restructured biased prompts effectively in most cases, providing a faster, more scalable guardrail, although it did not outperform human intervention.
Who Should Read "Managing LLM Bias in Investing"?
This report is intended for the following:
- Portfolio managers
- Investment analysts
- Chief investment officers
- Risk professionals
- Data scientists
- Model governance teams
- Compliance professionals
- Investment firms implementing generative or agentic AI
It should also be valuable to regulators, policymakers, consultants, and asset owners evaluating responsible AI practices within financial services.
What Will You Learn?
Recognize the sources of bias. Separate human bias, implicit model bias, and observable workflow decisions rather than treating the LLM as an isolated system.
Detect bias through comparison. Systematically vary framing, context, data order, and model choice, and then measure how outputs change across repeated tests.
Build stronger safeguards. Use balanced prompts, predefined sources and analytical steps, checkpoints, independent model review, and escalation thresholds for human review.
Create task-specific governance. Develop internal benchmarks for frequent or high-risk use cases and retest workflows as models, tools, and data change.
Why Is This Report Important Now?
As firms embed LLMs into higher-stakes investment workflows, bias can spread faster and become more difficult to identify. A plausible explanation or polished output is not evidence of impartial analysis. Firms need repeatable tests, documented controls, and clear accountability before relying on AI-supported conclusions.
How Can Investment Professionals Manage LLM Bias?
They can start with education and experimentation. Define the bias relevant to a specific task, test controlled prompt variants, quantify differences, and document acceptable thresholds. Mask company names and dates when appropriate, compare models from different providers, constrain data and tool selection, and require human review for consequential decisions. For recurring workflows, use structured prompt templates or a separate bias-detection model to screen inputs and outputs — but verify its judgments rather than treating the guardrail as infallible.
Frequently Asked Questions
What is LLM bias in investing?
LLM bias is a systematic tendency that can influence AI-generated investment analysis, recommendations, and decisions. It might arise from model training and architecture, from the information supplied by a user, or from the model’s choices about sources, tools, and analytical steps.
How can framing bias affect investment analysis?
Framing bias occurs when equivalent information produces different judgments because it is presented as a gain rather than a loss, or vice versa. In the report’s case study, positively and negatively framed versions of the same investment data led to different LLM evaluations, especially for drawdown performance and client retention.
Can prompt engineering eliminate LLM bias?
No. Prompt engineering can reduce particular biases, but it cannot remove every source of bias. Effective controls combine balanced prompts with repeated testing, documented data and tool choices, model comparison, bias-detection guardrails, and human review.
Which mitigation method worked best in the case study?
Manually restructuring prompts to include both positive and negative perspectives produced the largest reduction in framing bias. A separate bias-detection LLM also reduced bias in most tests and could provide a faster, more scalable screening step, but it did not outperform human intervention.
When should investment firms require human review?
Human review is most important when an AI-supported output might influence a consequential, high-risk, or difficult-to-reverse decision. Firms should define escalation thresholds, record who approves the output, and retain enough evidence to audit the complete human + AI workflow.
Key Takeaways
- Bias comes from both people and AI.
Investment outcomes can be influenced by human cognitive biases, the way prompts are written, model design, training data, and the decisions an AI system makes during a workflow. Managing bias requires looking at the entire human + AI process, not just the model. - How information is presented matters.
The research found that identical financial information can produce different LLM evaluations when it is framed positively or negatively. Drawdown performance and client retention showed the strongest and most consistent framing effects. - Prompt engineering alone is not enough.
Simply telling an LLM to avoid bias did not significantly reduce biased outputs, nor did defining framing bias within a prompt. More-structured approaches are needed. - Human oversight remains the most effective safeguard.
Manually restructuring prompts to present balanced positive and negative perspectives produced the greatest reduction in framing bias. A separate bias-detection LLM also improved results but did not outperform human review. - Bias can be managed through structured governance.
Firms should use comparative testing, balanced prompt templates, predefined analytical steps, internal benchmarks, model comparison, and clear escalation points for human review to make AI-supported investment decisions more reliable and auditable. - The goal is not bias-free AI — it is a trustworthy investment workflow.
Bias cannot be eliminated completely, but it can be identified, measured, and mitigated. Responsible AI adoption depends on transparent workflows, documented controls, continuous testing, and informed human judgment.