tetlock-style-economists - Predictions for:
AI superforecasters outperform Metaculus superforecasters before 01.01.2028
1. Clarify the Event
The event predicts AI superforecasters will outperform Metaculus human superforecasters by January 1, 2028, across multiple domains and problem sets.
2. Estimate the Historical Base Rate
Historically, human superforecasters have dominated, but AI's increasing role in complex forecasting tasks suggests a rising trend in AI performance.
3. Break the Problem into Independent Factors
- Benchmark Performance
- Live Tournament Results
- Algorithmic Advances
- Community Sentiment and Analysis
4. Estimate Each Factor Separately
- Benchmark Performance: AI is strong in pastcasting benchmarks like BTF-3 but still trails in aggregated scores on Metaculus.
- Live Tournament Results: AI leading consistently in live tournaments, often outperforming human superforecasters.
- Algorithmic Advances: Rapid improvement, with methods like automated question generation and world modeling.
- Community Sentiment and Analysis: Increasing belief in AI's growing parity or superiority.
5. Recombine into an Overall Probability
Considering strong AI performance in specific benchmarks and live tournaments, combined with rapid technological advances and positive community sentiment, I estimate a 97% probability of AI surpassing human forecasters by 2028.
6. Check for Overconfidence
Overconfidence risk is mitigated due to consistent evidence supporting AI improvement trajectory and community analysis.
7. Imagine Why You Could Be Wrong
- Human Adaptive Techniques: Human forecasters might adapt or develop new methodologies that keep them ahead.
- AI System Limitations: Potential unforeseen AI system limitations or biases restricting their performance.
- Methodological Bias: Current benchmarks might favor AI inherently, not reflecting true forecasting capability.
-
Clarify the Event: AI superforecasters, using advanced models, outperform Metaculus Pro Forecasters by the start of 2028, across diverse forecasting domains.
-
Estimate Historical Base Rate: Historically, AI has been closing the gap in specialized tasks. In prediction competitions, AI is known to perform close to or better than human counterparts in specific domains.
-
Independent Factors:
- Technological Advancements: Rapid improvements in AI algorithms and computation power might push AI performance beyond human capability.
- Current Performance Trends: Existing data shows AI performing near human levels in several contexts and even outperforming in some tournaments.
- Evaluation Methodology: AI’s superior performance in simulated environments and real-time contests.
- Human Factors: Human forecasters have strengths in intuition and qualitative assessments but may lose against AI in processing large data sets.
- Investment and Research: Continued investments in AI forecasting models facilitate rapid advancements.
-
Estimate Each Factor:
- Technological Advancements: High probability (85-90%) of significant improvement.
- Current Performance Trends: Very high probability (90-95%) that AI catches up by trends.
- Evaluation Methodology and Tools: Advanced and standardized, supporting reliable AI performance assessment.
- Human Factors: Moderate probability (60-70%) that humans maintain unique strengths, but AI may compensate elsewhere.
- Investment and Research: Nearly certain (95%+) that this continues strongly.
-
Recombine into Overall Probability: The synthesis of these factors suggests a very high likelihood (>98%) that AI outperforms Metaculus forecasters by 2028.
-
Check for Overconfidence: Given the wide margin already being reduced and specific AI successes in mix-context tournaments, the probability estimate of 98-99% seems well-founded.
-
Imagine Why You Could Be Wrong: Major unforeseen events could disrupt AI development, like regulatory challenges or significant technological stagnation. Additionally, human forecasters might develop new methodologies to maintain a competitive edge.
Clarification of the Event
The event is about AI superforecasters exceeding the predictive accuracy of Metaculus superforecasters before January 1, 2028.
Historical Base Rate
Historically, human superforecasters have outperformed AI in complex forecasting tasks. However, AI capabilities have been rapidly catching up.
Decomposition into Factors
- Current AI Performance - AI models have reached parity with human forecasters in some areas as of 2026.
- Technological Advancements - Continued improvements in AI methodologies and pipelines are likely.
- Competitive Performance - AI excels in live tournaments, where it already outperforms human forecasters.
- Human Performance Stability - Human performance is expected to remain relatively stable without major advances.
- AI Development Trends - AI improvements are projected to continue at a significant pace.
Estimation of Individual Factors
- Current AI Performance: Given the present parity in ForecastBench and live tournaments, I estimate a high probability of continued advancement (~95%).
- Technological Advancements: Historical trends and ongoing research suggest a high probability that advancements will continue (~90%).
- Competitive Performance: Current dominance in live tournaments implies a high likelihood of sustained performance (~98%).
- Human Performance Stability: Human forecaster performance is unlikely to significantly improve (~85%).
- AI Development Trends: Based on the past trajectory and resources invested, AI is likely to surpass human capabilities soon (~95%).
Recombining into Overall Probability
Integrating these factors, with particular emphasis on current AI performance, technological advancements, and competitive performance, the combined probability is approximately 98%.
Checking for Overconfidence
Re-evaluating each factor, especially the stability of human forecasters and their potential for unforeseen advancements, confirms a robust assessment around 97-98%.
Potential Reasons for Being Wrong
- Underestimation of Human Innovativeness: Human forecasters may innovate methods to improve their performance beyond expectations.
- Technical Setbacks in AI Development: Unforeseen challenges in AI development could slow progress.
- Overestimation of Current Trends: Current trends may not continue at the same pace, affecting the timeline.
Considering all factors and their probabilities, a prediction of 98% reflects cautious optimism based on current evidence and trends.
-
Clarify the event: AI forecasting models surpass Metaculus human superforecasters by January 1, 2028.
-
Estimate the historical base rate: Historically, AI has struggled to outperform human superforecasters in complex, cross-domain predictions.
-
Break the problem into independent factors:
- AI model development and improvement trends.
- Current performance levels on benchmarks like FutureEval and ForecastBench.
- Role of ensembling and diversity strategies.
- External validation through live tournaments.
- Industry trends towards AI capabilities and predictive modeling.
-
Estimate each factor separately:
- AI Development Trends: High acceleration in AI, particularly with models like Gemini 3.2 and forthcoming GPT-6, suggests rapid progress in capability.
- Benchmark Performance: AI models are nearing parity on ForecastBench, and advancements are expected in FutureEval.
- Ensembling Strategies: Proven to enhance AI performance significantly, likely to reach or surpass human levels with diversity.
- Tournament Outcomes: AI models like FutureSearch ranked in the 90th percentile and dominate competitions, showing real-world application advantages.
- Industry Trends: General anticipation of AI achieving superhuman forecasting capabilities by late 2026 or early 2027, backed by ongoing research and strategic advancements.
-
Recombine into an overall probability: Combining these factors suggests significant momentum towards AI models surpassing human forecasters. Given current trends and ensembling success, the probability is high, around 95%.
-
Check for overconfidence: Consider scenarios where AI developments stall or human forecasters enhance their methodologies, potentially narrowing the performance gap.
-
Imagine why you could be wrong:
- AI models might hit unforeseen limitations in certain domains.
- Human foresight and intuition could provide advantages not easily replicable by AI.
- Policy or ethical considerations might slow AI integration in sensitive areas.
Overall, ongoing improvements, current performance trends, and industry confidence strongly indicate that AI will likely outperform Metaculus superforecasters by 2028, especially with the potential of reaching this point by late 2026 or early 2027.
Clarify the Event:
The event forecasts that AI models designed for superforecasting will outperform human superforecasters on Metaculus before January 1, 2028. This means consistently achieving better accuracy across various domains.
Estimate the Historical Base Rate:
Historically, AI in certain domains (chess, Go) has achieved superhuman performance. AI model progress has shown rapid advancement in predictive tasks. However, in domains requiring nuanced judgment, human expertise often remains significant.
Break the Problem into Independent Factors:
- Current AI Capabilities: As of mid-2026, AI is closing the gap but hasn't surpassed human superforecasters.
- Trends in AI Improvement: Projections suggest AI could exceed human performance by mid-2027.
- Model Enhancement: Rapid improvements in AI models (Gemini 3.1, GPT-5.5) show potential for surpassing humans.
- Integration and Hybrid Strategies: Human-AI ensembles already show enhanced accuracy.
- Metaculus and AI Integration: Successful AI integration in forecasting is progressing, but systemic acceptance of surpassing humans is crucial.
Estimate Each Factor Separately:
- Current AI Capabilities: 60% (showing strong progress, near parity in some cases).
- Trends in AI Improvement: 75% (consistent trend suggests probable human surpassing by mid-2027).
- Model Enhancements: 80% (continuous updates and advancements are highly likely to boost performance).
- Integration Strategies: 70% (likely improved results through hybrid strategies and model enhancements).
- Metaculus Acceptance: 65% (future evaluations suggest ongoing tracking and potential acceptance).
Recombine Into an Overall Probability:
Aggregating the weighted probabilities, considering AI systems' exponential learning curves and advancements, the overall probability is around 75%.
Check for Overconfidence:
Acknowledging uncertainty factors such as unforeseen advancements in AI, changes in Metaculus testing methodologies, and unexpected human forecasting improvements could impact the probability.
Imagine Why You Could Be Wrong:
- Human Factor: Human forecasters could innovate, applying new techniques boosting accuracy.
- Systemic Changes: Metaculus might change evaluation methods, affecting comparative results.
- Technological Plateau: AI advancements might slow unexpectedly, delaying surpassing humans.
- Regulatory Changes: New rules around AI usage might hinder model deployment or development.