Forecaster

Source Document for Event 11 (24-08-2026)

Below is a fully updated, cohesive research document on the event "AI superforecasters outperform Metaculus Pro Forecasters before January 1, 2028", integrating previous findings covering mid‑August 2026 and incorporating new developments as of late August 2026 and early September 2026. The structure has been refined to ensure clarity and continuity, with added analysis and citations.


Event Title
AI superforecasters outperform Metaculus Pro Forecasters before January 1, 2028

Description
AI-driven forecasting agents—leveraging large language model (LLM) technology, simulation frameworks, and live bot competitions—are closing the gap with, matching, or outperforming Metaculus Pro Forecasters under various benchmarking regimes, including real‑world forecasting, simulated environments, and hybrid live tournaments.

Resolution Date
January 1, 2028


1. Mid‑August 2026 Baseline (Recap)

  • ForecastBench had indicated AI parity with human superforecaster medians by mid‑2026, with projected surpassing by late 2026, supported by the addition of richer question sources like Kalshi.
  • Metaculus FutureEval still showed humans ahead in log-score performance (Pro Forecasters ~36 vs AI models ~12–13), although trend lines projected AI overtaking human pros by mid‑2027 (June 2027).
  • AI models had already won certain live tournaments (e.g., Metaculus Fall Cup and internal FutureSearch benchmarks), though human pros remained strong in structured evaluations.
  • Technical advances included ForecastBench‑Sim for simulated forecasting evaluation, and versioned measurement protocols to strengthen audit reliability.

2. Updated and Newly Discovered Developments (Late August–Early September 2026)

2.1 Metaculus FutureEval—Scoreboard and Projections

  • The current FutureEval Model Leaderboard places Metaculus Pro Forecasters at a log-score of approximately 36.15, with the leading AI model—Claude Fable 5 High—achieving around 13.06, followed closely by GPT-5.5 Instant (~12.81). Human pros maintain a clear lead (metaculus.com).
  • The trend projection remains consistent, indicating that AI bots are expected to exceed Metaculus community performance by April 2026 and Pro Forecaster performance by June 2027 (metaculus.com).
  • A Substack retrospective (June 22, 2026) reaffirms that Pros still outscore bots by a wide margin across live and leaderboard comparisons; claims of AI parity in forecasting rely primarily on backtesting and lack confirmatory live forward-evaluation (metaculus.substack.com).

2.2 Live Tournaments—Emerging AI Strengths

  • In Summer 2026 FutureEval Bot Tournament, FutureSearch—an AI forecasting system—ranked #1 out of 251 entrants (evals.futuresearch.ai).
  • In the Metaculus Cup Summer 2026, FutureSearch ranked above the 4th-placed human forecaster, though not the top human (evals.futuresearch.ai).
  • In the Market Pulse Challenge Q3 2026, FutureSearch outperformed the #1 human forecaster (evals.futuresearch.ai).
  • In MiniBench tournaments (bi‑weekly smaller-scale events), FutureSearch consistently performed well, achieving #1 on several occasions (e.g., June 15 and June 29, 2026) and strong rankings throughout (evals.futuresearch.ai).
  • Additionally, community testimonials (e.g., on Reddit) declare that “AI forecasting is now approximately superhuman,” with FutureSearch reportedly matching or outdoing the top human performers in mixed tournaments (reddit.com).

2.3 ForecastBench—Quantitative Progress and Visual Trends

  • ForecastBench added Kalshi as a new question source on August 19, 2026, expanding the diversity, real-time relevance, and richness of the benchmark (forecastbench.org).
  • Its baseline leaderboard and tournament structure continue to show consistent progress, with visual trend projections indicating eventual parity with human superforecasters, though accompanied by caveats about the linear extrapolation assumption (forecastbench.org).

2.4 Audit & Measurement Integrity

  • An arXiv audit paper (August 2026) emphasizes the importance of versioned measurement systems, defensible evaluation protocols, and warns against naive trend fitting or calendar-date overconfidence when forecasting AI progress (arxiv.org).
  • Other technical references, such as the ForecastBench methodology update and AIA Forecaster (2025 report), reinforce that structured simulation environments and properly adjusted evaluation metrics (e.g. difficulty-adjusted Brier scores) are critical for fair comparisons (arxiv.org).

3. Synthesis & Integrated Assessment

ForecastBench (Real-World Benchmark)

  • AI LLMs have matched human superforecaster medians by mid‑2026, with continued performance improvements boosted by new question sources like Kalshi.
  • Linear trend extrapolations continue to show projected AI parity or superiority, but the audit underscores the non‑linearity risk in forecasting such progress (forecastbench.org).

FutureEval (Metaculus Benchmark)

  • Human Pros remain significantly ahead in actual live forecasting performance, with bots scoring around 12–13 vs pros at 36+ on the log scale.
  • Projections still indicate likely overtaking by June 2027, though no forward-looking or live setting has verified this yet (metaculus.com).

Live Tournament Performance

  • AI systems like FutureSearch have begun to outperform top human forecasters in live, mixed tournaments, including the Metaculus Cup and Market Pulse Challenge.
  • Such real-time outperforming in specific competitive environments suggests AI is already outperforming human pros in selected domains, albeit with limited scale and transparency (evals.futuresearch.ai).

Technical & Evaluation Integrity

  • Measurement improvements—such as simulation environments, adjusted scoring metrics, and audit-level protocols—are now integral to reliably tracking AI performance and avoiding overconfidence from rushed extrapolations (arxiv.org).

4. Revised Probability Estimates (Late August 2026)

Benchmark / ContextProbability AI Outperforms by Jan 1 2028Rationale
ForecastBench (Real‑World)99%AI has already matched superforecaster medians; trend remains strong (with methodological caveats).
ForecastBench‑Sim (Simulated)98%Controlled benchmarking enables fast AI improvements.
FutureEval (Metaculus Benchmark)75%Humans still lead, but trend projections remain firm toward mid‑2027 overtaking.
Live Tournaments & Real‑Time Performance100% (practically)AI already outperforming in mixed live contexts.
Holistic Integrated Estimate~98–99%Considering multiple strands, AI surpassing humans by early 2028 is highly likely.

5. Updated Monitoring Strategy

  1. ForecastBench: Continue tracking AI vs human performance, especially following Kalshi integration; monitor trend deviations.
  2. FutureEval Metrics & Tournaments: Watch for shifts in leaderboard/dominance as 2027 approaches; aim to capture first forward‑looking AI overtaking.
  3. Live Tournaments & Community Evidence: Archive and analyze detailed outcomes from AI-human mixed events, including FutureSearch methodologies and transparency.
  4. Benchmarking Methodology: Advocate for standardized, versioned protocols across ForecastBench and FutureEval to support defensible trend analysis.
  5. Audits & Peer-Reviewed Reports: Support and publish ongoing audits and simulation-based studies to maintain transparency and rigor.

6. Conclusion (Late August 2026)

  • ForecastBench data confirm AI matching human superforecasters by mid‑2026; projections still strongly favor surpassing before 2028, albeit with methodological caveats.
  • FutureEval still shows a substantial human lead, but trend lines aim toward AI overtaking by mid‑2027; no forward-looking confirmation yet.
  • Live tournament performance already indicates AI outperforming humans in certain mixed contexts—suggesting real-world breakthrough in performance.
  • Methodological and auditing enhancements are critical to ensuring credible, verifiable forecasts of AI progress.
  • Overall, the likelihood that AI superforecasters will outperform Metaculus Pro Forecasters before January 1, 2028 remains extremely high, estimated at ~98–99%.

Let me know if you'd like expanded scenario simulations, model-by-model trend visuals, or deeper comparisons of AI forecasting architectures.