← Back to list

The Enduring Tension between Human Intuition and Systemic Prediction

I’ve spent a good part of this weekend with a fascinating new paper on my screen, “The Strategic Foresight of LLMs: Evidence from a Fully…

Alireza Hejazi · 2026-02-08 09:19 · 0 claps · 4.3 min read
#ai #strategic-foresight #decision-making #predictions #humans
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General

The Enduring Tension between Human Intuition and Systemic Prediction

I’ve spent a good part of this weekend with a fascinating new paper on my screen, “The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament” (Csaszar et al., 2026). I was struck by how directly Csaszar and his colleagues speak to a core concern of our field: the enduring tension between human intuition and systemic prediction. Reading their paper felt like watching a significant data point land on a chart of a future we’ve been sketching for years. In the quiet of a Sunday afternoon, here’s a personal reflection on how this research sits with me.

The study’s contribution to the Global Foresight Community is foundational and timely. It directly tests the ability of Large Language Models (LLMs) to perform a quintessential foresight task in a methodologically rigorous, fully prospective setting.

Using live Kickstarter campaigns launched after the models’ training data cutoff, the authors created a prediction tournament where models and humans had to rank ventures by their future fundraising success. The results are humbling and consequential.

While experienced managers and MBA-trained investors achieved rank correlations with actual outcomes between 0.04 and 0.45, several frontier LLMs, led by Gemini 2.5 Pro (with a correlation of 0.74), demonstrated a clear and substantial superiority. In practical terms, the best AI correctly ordered nearly four of every five venture pairs.

For our community, this is more than just a “machine beats human” headline. It offers robust evidence that a new, scalable tool for augmenting and, in some cases, displacing one of the most prized human capabilities (judgment under deep uncertainty) is already operational.

From my perspective, the finding that neither aggregated human “wisdom-of-the-crowd” nor human-AI hybrid teams could beat the best standalone model is the most provocative insight. It suggests a paradigm shift rather than a simple tool upgrade.

The paper’s greatest strengths are its rigorous experimental design and its grounding in real-world, irreversible time. The authors’ choice of a fully prospective tournament elegantly sidesteps the critical pitfall of data leakage. In this way, they ensured that they were testing a genuine prediction, not just pattern retrieval.

By structuring the task as a series of 870 pairwise comparisons, they cleverly transform a complex ranking problem into a digestible format for both humans and LLMs, yielding stable and comparable results. Besides, their benchmark against a sizable cohort of experienced managers adds crucial context. It showed that the AI wasn’t just beating a low bar but outperforming credible professional judgment. The research doesn’t shy away from the theoretical implications. It thoughtfully situates its findings within the literature on bounded rationality and “unbounding rationality.”

No study is without its limits, and this one has a few important boundaries. The context, while excellent, is specific: Kickstarter campaigns have defined start and end points, clear success metrics (funds raised), and publicly available project narratives. The paper successfully shows that LLMs excel at distilling signals from such open-ended text to predict market reactions.

However, the most wicked strategic problems often lack this clarity; they involve shifting goalposts, multiple competing stakeholders, and outcomes that are not easily quantifiable. The study measures predictive accuracy, but strategic foresight also requires causal reasoning. We need to explain why one venture might succeed to inform action and adaptation.

As tested, the LLM is a powerful oracle, but not yet a strategist that can navigate the feedback loops and endogenous change it mentions in the introduction. While the human benchmark is solid, the sample of three expert investors is small. It leaves open questions about the absolute ceiling of unaided human expertise in this precise task.

My key takeaways from this research are threefold. First, we have crossed a threshold where AI can serve as a high-performance base layer for specific types of evaluative foresight, particularly those involving the synthesis of textual information to predict market outcomes.

Second, the failure of hybrid teams to outperform the best AI suggests that naive “human-in-the-loop” approaches may be insufficient. The value-add will come from framing the problems, curating the inputs, and, most importantly, deciding how to act on the predictions.

Third, the study provides a powerful methodological template. The prospective tournament is a tool the foresight community should adopt and adapt for benchmarking both human and machine capabilities in other domains of strategic uncertainty.

For growth and improvement, future research has a thrilling path ahead. I think that the immediate next step is to test this capability in messier, more endogenous strategic environments like competitive dynamics or long-term policy impacts.

I am particularly keen to see a work that moves from prediction to prescription. Specifically, I’d like to know: Provided with a prediction, can LLMs also generate and evaluate potential interventions to alter the forecasted outcome?

Another critical avenue is to unpack the “black box.” I’d like to know: What specific cues in the project narratives did the LLMs leverage that humans overlooked or undervalued? Understanding this could train better human judgment even as it improves AI.

In my view, exploring how to effectively couple human strategic intuition with AI’s predictive power to create a true “centaur” system (one that outperforms either component alone) remains the essential, unsolved challenge.

This paper doesn’t spell the end of human strategists; it recalibrates our role. The future of foresight may well lie in moving “upstream” to define the problems and “downstream” to implement the insights, while leveraging tools like these for the computationally intense prediction in the middle.

It was a worthwhile weekend read, leaving me both impressed by the technical achievement and thoughtfully unsettled about the implications, which is exactly what good futures research should do. I offer my best gratitude to the authors for co-authoring such an informative piece. Thank you so much!

Reference

Csaszar, F. A., Peterson, A., & Wilde, D. (2026). The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2602.01684

© Alireza Hejazi 2026

Alireza Hejazi, Ph.D., is an independent futures researcher and leadership consultant with combined corporate and academic experience. He has developed several assessment instruments, including the Corporate Foresight Index and other evaluative tools for foresight and leadership. Alireza is the author of 10 books on leadership and foresight, such as *Cognitively Adaptive Leaders, [Foresight Evaluation Frameworks](https://wp.me/pc7Da5-1Bi), [Foresight in Flux](https://wp.me/pc7Da5-1zy), The Strategic Foresight Workshop Handbook, The Art of Selling Foresight, Responsible Foresight, and Becoming A Professional Futurist*.


메타데이터
post_id
3cf46316cf58
slug
the-enduring-tension-between-human-intuition-and-systemic-prediction-3cf46316cf58
url
https://medium.com/@hejazi/the-enduring-tension-between-human-intuition-and-systemic-prediction-3cf46316cf58
canonical_url
https://medium.com/@hejazi/the-enduring-tension-between-human-intuition-and-systemic-prediction-3cf46316cf58
author_url
https://medium.com/@hejazi
status
ok
fetched_at
2026-06-11 12:34:08