Forecaster

Source Document for Event 11 (09-08-2026)

Below is an initial, comprehensive research document for the event forecasting task:

Event Title:
AI superforecasters outperform Metaculus superforecasters before January 1, 2028

Description:
AI-based forecasting agents (bots/models) consistently achieve better forecasting accuracy (e.g., via log score or Brier score) than Metaculus’s top human forecasters (Pro or "superforecasters") across diverse domains and question sets, up to the resolution date.

Resolution Date: January 1, 2028


1. Background and Context

1.1 Superforecasting and Metaculus

  • Superforecasters are individuals who consistently outperform others in forecasting accuracy over time, as popularized by Tetlock and Mellers. Their rarity and robustness were demonstrated in both controlled experiments and real-world projects (sciencedirect.com).
  • Metaculus is a public forecasting platform aggregating predictions, including contributions from both a broad community and a smaller group of high-performing forecasters. It also recently launched the “FutureEval” benchmark to measure AI vs. human forecasting accuracy (en.wikipedia.org).

1.2 AI Forecasting on Metaculus (FutureEval)

  • FutureEval, launched in February 2026, tracks AI model performance on real-world forecasting questions and compares it to both the aggregate Metaculus community and its Pro forecasters. Projections from Metaculus indicated that AI might surpass community-level performance by April 2026 and Pro Forecasters by mid‑2027 (globenewswire.com).
  • The platform’s Model Leaderboard continuously updates AI model performance using standardized prompts and scoring methods, enabling trend visualization and forecasting of when AI may reach or exceed human baselines (metaculus.com).

1.3 Recent Performance Trends

  • As of mid‑2026, AI “bots” on Metaculus remain behind Pro Forecasters, though trends show rapid improvement (metaculus.com).
  • In Q4 2024’s AI benchmarking tournament, bots narrowed the performance gap but still trailed the human Pro group, though without statistical significance (p = 0.079) (lesswrong.com).
  • Independent benchmarks, such as ForecastBench, suggest AI models may already be statistically indistinguishable from human superforecasters in some contexts (forecastingresearch.substack.com).

2. Empirical Evidence and Research Insights

2.1 Academic Evaluations

  • A mid‑2025 study evaluating LLMs on 464 Metaculus questions found that frontier models surpassed the general human crowd but still underperformed compared to human superforecasters (arxiv.org).
  • Another study in early 2024 showed that LLMs could augment human forecasting accuracy by 24–28%, indicating their utility in improving human performance (arxiv.org).
  • Prior experiments with GPT-4 in a three-month Metaculus tournament (Jul–Oct 2023) found its forecasts significantly worse than the median human crowd; it did not outperform even a naïve 50% forecasting strategy (arxiv.org).
  • Yet, ensemble strategies (aggregating LLM predictions) have shown parity with human crowds and improvement when combined with human forecasts (pmc.ncbi.nlm.nih.gov).

2.2 Automated Question Generation & Tournaments

  • A 2026 ICLR‑presented system automatically produced and resolved diverse forecasting questions using LLMs, achieving high-quality generation and using resulting data to evaluate how LLMs like Gemini 3 Pro and GPT‑5 performed (e.g., Brier scores around 0.13–0.18) (arxiv.org).

3. Leading Indicators and Trends Toward the Event

3.1 Trend Data from FutureEval

  • Metaculus’s own projections (as of Feb 2026) indicate AI models could outperform the community by April 2026 and surpass Pro Forecasters by mid‑2027 if current trends continue (globenewswire.com).

3.2 Tournament Results

  • Q4 2024 bot performance was significantly closing the gap: head-to-head score at –8.9 down from –11.3 in Q3, signaling improvement toward parity (lesswrong.com).

3.3 Independent Benchmarking

  • ForecastBench analysis in mid-July 2026 indicates several AI models are already statistically indistinguishable from human superforecaster accuracy (forecastingresearch.substack.com).

3.4 Ensemble Advantages

  • Combining multiple LLMs (silicon crowd) yields performance that matches human crowds, and hybrid human-AI forecasts further enhance accuracy by 17–28% (pmc.ncbi.nlm.nih.gov).

3.5 Rapid Model Progress

  • Metaculus Model Leaderboard shows recent high-performing frontier models (Gemini 3.1 Pro, GPT‑5.5 variants, Claude Opus) with clear numeric skill scores, reflecting rapid advancement (metaculus.com).

4. Summary of Current Evidence

At present (mid‑2026):

  • Bots consistently outperform the broader Metaculus community, but still lag behind the Pro forecasters.
  • Bots are rapidly closing the gap: predictive performance improving quarter-over-quarter.
  • ForecastBench and ensemble LLM methods indicate AI can reach parity with human superforecaster accuracy in some contexts.
  • Metaculus projections suggest AI could surpass Pro Forecasters by mid‑2027, ahead of the January 1, 2028 resolution date.

5. Key Variables to Monitor Going Forward

To support forecasting agents evaluating this event, the following variables and indicators are critical:

  1. FutureEval Leaderboard trends – continuous monitoring of AI vs Pro Forecasters skill scores and projected crossover points.
  2. Quarterly Tournament Results – particularly head-to-head comparisons and p-values to assess when AI surpasses humans with statistical significance.
  3. ForecastBench Updates – ongoing independent evaluation of AI vs superforecasters across diverse questions.
  4. Ensemble & Hybrid Strategies – performance of model ensembles and human-AI aggregations.
  5. Advances in Frontier Models – tracking model releases (Gemini 3.2 Pro, GPT‑6, Claude Opus 5.0, etc.) and their FutureEval performance.
  6. Research Publications – new academic results evaluating model forecasting on Metaculus or similar platforms.
  7. Tournament Structure Changes – e.g., shifts in question types, scoring, or bot participation rules affecting performance comparability.

6. Preliminary Assessment (as of August 2026)

  • Current evidence suggests the event is plausible and likely: AI forecasters are trending toward outperforming Metaculus superforecasters, with mid‑2027 projected crossover.
  • Independent benchmarks show models already approaching human superforecaster performance.
  • Thus, barring unexpected slowdowns or systemic changes, the event occurring before January 1, 2028 remains a credible outcome.

7. Research Document Summary

  • Event: AI outperforms Metaculus superforecasters by Jan 1, 2028.
  • Status: AI lags but closing fast; crossover projected mid‑2027.
  • Evidence: FutureEval projections, tournament data, independent benchmarking, ensemble gains.
  • Indicators to Watch: Leaderboard trends, tournaments, model advances, research outputs.
  • Initial Forecast: Probability increasing; close to realistic 50–70% range depending on trend continuation.

This completes the initial research summary. It provides a structured, well-cited foundation for forecasting agents who will assess the probability of the specified event.

Let me know if you would like deeper dives into specific areas (e.g., individual model performance, statistical significance timelines, Trend analysis charts, or ensemble strategies).