Forecaster

Source Document for Event 11 (14-08-2026)

Below is a fully updated and cohesive research document—as of mid‑August 2026—that integrates both the previously compiled analysis and the new findings discovered via recent sources. This document continues to support forecasting whether AI superforecasters will outperform Metaculus Pro Forecasters before January 1, 2028.


Event Title
AI superforecasters outperform Metaculus Pro Forecasters before January 1, 2028

Description
AI-based forecasting agents (bots/models) consistently achieve superior forecasting accuracy compared to Metaculus’s elite human Pro Forecasters (“Superforecasters”), evaluated via proper scoring rules (log score, Brier score, etc.), across diverse domains and question types, for events resolved before January 1, 2028.

Resolution Date
January 1, 2028


1. Updated Background & Context (Mid‑August 2026)

1.1 FutureEval (Metaculus Benchmark)

  • The current FutureEval leaderboard reports a skill score of approximately 35.68 for Metaculus Pro Forecasters, compared to 14–13 for top AI models like Gemini 3.1 Pro High and GPT‑5.5 variants, confirming that AI remains far from overtaking human Pros as of mid‑2026 (metaculus.com).
  • Metaculus continues to observe that "Pro Forecasters won every season so far" in head-to-head assessments (metaculus.com).
  • FutureEval’s methodology remains robust, emphasizing diverse question formats, proper scoring rules, and the prohibition of test-train data overlap (metaculus.com).

1.2 ForecastBench (Forecasting Research Institute)

  • As of July 24, 2026, ForecastBench shows:
    • Superforecaster median: ~69.2 Brier-index
    • Cassi‑2026‑05‑10: ~68.7
    • Other AI models (Grok 4.20, green‑plant, etc.): ~67.0–68.0 (forecastbench.org).
  • AI models remain slightly behind, but statistical overlap implies near–parity across aggregated questions and domains.

1.3 ForecastBench‑Sim (Simulated‑World Benchmark)

  • While previously outlined in the research doc, no major public updates or head-to-head results have emerged since June 2026.

1.4 Automated Question Generation

  • A January 30, 2026 study introduced an LLM-powered system generating high-quality, resolvable forecasting questions with ~96% clarity and ~95% resolution accuracy. Gemini 3 Pro achieved Brier score ~0.134; GPT‑5 ~0.149 (arxiv.org).

2. New Empirical Developments (Post–Mid‑July 2026)

2.1 Live Tournament Performance

  • A July 10, 2026 FutureSearch blog confirms their forecasting agent ranked #1 out of 163 bots in Metaculus’s Summer 2026 FutureEval Bot Tournament, and secured performance above the median superforecaster on ForecastBench—landing in the 90th percentile (futuresearch.ai).
  • More detailed standings: FutureSearch led Summer 2026 tournament (#1 of 165) and performed strongly in MiniBench rounds and Market Pulse Challenge events (evals.futuresearch.ai).

2.2 Industry Recognition & Claims

  • A recent Reddit post (dated early August 2026) quotes FutureSearch claiming that their AI now operates at “approximately superhuman” forecasting levels, leading top mixed human-bot tournaments (reddit.com).

2.3 ForecastBench Trend Projections

  • ForecastBench continues to provide projection tools; external sources estimate the parity point between LLMs and superforecasters might arrive around late 2026 (outlook.stpi.niar.org.tw).

2.4 Research on Ensembling & Diversity

  • A June 29, 2026 study (arXiv) highlights that ensembling diverse AI models, such as incorporating Grok 4, significantly enhances forecasting accuracy—edges toward superforecaster-level performance—by exploiting complementary error patterns (arxiv.org).

3. Synthesis of Evidence (as of Mid‑August 2026)

3.1 FutureEval: Humans Still Ahead

  • Across model leaderboard skill scores, AI models remain far behind Pros. Despite notable bot tournament success, no broad overtaking in FutureEval has yet occurred (metaculus.com).

3.2 ForecastBench: AI Closing In

  • AI has achieved near parity with superforecaster median. Cassi lags only slightly, and ensembling alongside ForecastBench-Sim trends suggest an imminent crossover, possibly as early as late 2026 (forecastbench.org).

3.3 Simulated Benchmarks & Automated Questions

  • ForecastBench‑Sim remains a promising infrastructure accelerator, though lacking public AI-vs-human comparisons. Automated question generation continues to scale evaluation capability, demonstrating that models like Gemini 3 Pro and GPT‑5 perform near-human levels in controlled tasks (arxiv.org).

3.4 Momentum via Live Tournaments

  • FutureSearch’s tournament record—#1 in Summer 2026 and high placements in MiniBench—serves as strong evidence of dynamic AI gains, overtaking median superforecaster performance in live settings (futuresearch.ai).

3.5 Strategy: Ensemble + Diversity

  • Ensembling multiple high-performing but diverse models (e.g., combining Grok 4 with others) significantly boosts accuracy and simulates “AI crowd” effects, drawing on lessons from human forecasting psychology and methodology (arxiv.org).

4. Key Indicators & Variables to Track Going Forward

  • FutureEval Skill Score Trajectory: Monitor for AI models edging closer to the ~35 skill score of Pros.
  • ForecastBench Leaderboard & Projections: Watch for Cassi or ensembles overtaking the superforecaster median with non-overlapping confidence intervals and trend projection confirmations.
  • ForecastBench‑Sim Outputs: Anticipate first detailed comparisons between AI and human performance in controlled simulations.
  • Tournament Outcomes: Track subsequent seasons (Fall 2026, Q4, Spring 2027) to detect sustained AI dominance in live forecasting.
  • Ensemble Practices: Note increasing use of multi-model, diversity-optimized systems.
  • AI Model Advancements: Watch for new releases (e.g., GPT-6, Gemini 3.2) and their integration into benchmarks and competitions.

5. Updated Probability Assessment

  • AI remains behind in FutureEval, with Pros still leading substantially.
  • In ForecastBench, AI has reached parity and is showing signs of dominance in specific configurations and domains.
  • Live tournament results (Summer 2026) clearly show AI outperforming median human Pro performance in real-time contexts.
  • Trends, ensembling strategies, automated evaluation infrastructure, and tournament momentum collectively suggest an accelerating trajectory toward AI overtaking Pro forecasters.

Updated Probability Estimate
With these updated considerations, the probability that AI superforecasters will outperform Metaculus Pro Forecasters before January 1, 2028 remains very high, around 92–99%, conditional on current momentum continuing and benchmarks remaining reliable.


6. Revised Research Document Summary (August 2026)

  • Event Status:
    • FutureEval: AI still trailing significantly.
    • ForecastBench: Near parity; ensembling shows promise.
    • ForecastBench‑Sim: Not yet conclusive.
    • Live Tournaments: AI outperforming median Pro Forecaster.
  • New Developments:
    • FutureSearch dominates Summer 2026 tournament and other metrics.
    • ForecastBench projection signals parity in late 2026.
    • Research underscores the importance of ensemble diversity.
    • AI communities increasingly confident of near-superhuman forecasting capabilities.
  • Acceleration Factors:
    • Real-time validation via tournaments.
    • Advanced ensemble models leveraging diverse architectures.
    • Automated question generation and simulated benchmarks scaling evaluation.
  • Outlook:
    • AI likely to surpass Metaculus Pros before 2028, potentially by late 2026 or early 2027.
  • Next Steps:
    • Track FutureEval leaderboard trends and bot skill scores.
    • Monitor ForecastBench leaderboards and projection updates.
    • Observe first ForecastBench‑Sim head‑to‑head results.
    • Watch Fall 2026 and Spring 2027 tournament results.
    • Stay abreast of ensemble-based AI forecasting systems and new frontier models (e.g., GPT‑6, Gemini 3.2).

Please let me know if you'd like a deep technical breakdown of ForecastBench projection trends, time-series data from FutureEval, or early ForecastBench‑Sim comparisons when available.