The Future is
the Label.

Peer-reviewed research. Live benchmark wins. Government-cleared.

Jul 2026Launch

Foresight v4 sets the Pareto frontier for AI forecasting

We released Foresight v4: a new generation of forecasting models that push the accuracy-cost frontier, with low-cost and full-strength configurations for production forecasting agents.

Brier Skill Score on resolved Polymarket questions
higher is better · cheaper is left
Foresight Frontier models
Foresight v4 Pareto frontierScatterplot showing Foresight v4 Low and Full above frontier AI models on Brier Skill Score for comparable forecast cost.Brier Skill Score26%19%13%GPT-5.4GPT-5Opus 4.6Gemini 3.1 ProForesight v3Foresight v4(Low)Foresight v4(Full)
all-in cost per 1,000 forecasts →
Jul 2026Benchmark

Foresight v3 reaches superforecaster range on ForecastBench

Foresight v3 scores within range of the superforecaster median on ForecastBench, with no statistical superforecaster edge. Foresight v3 beats every other frontier AI entry from OpenAI, Anthropic, and xAI, with only Gemini marginally ahead—while running on a single GPU.

ForecastBench leaderboard showing Foresight v3 within range of the superforecaster median, marginally behind Gemini, and ahead of GPT-5 and Grok entries
Jul 2026Spotlight

Future-as-Label earns spotlight in AI Forecasting @ ICML

Future-as-Label, our core proprietary methodology for training AI from real-world outcomes, was selected as one of ten spotlights at the ICML 2026 Workshop on AI Forecasting.

ICML logo
ICML 2026
AI Forecasting Workshop
Spotlight
Workshop paper
May 12, 2026Research

Training Large Language Models to Predict Clinical Events

Foresight Learning trains AI to predict clinical events directly from raw clinical notes, with no hand-labeled dataset. It learns from outcomes that appear later in the patient record, letting hospitals build calibrated predictors for their own patient populations. On MIMIC-III, our GPT-OSS-120B model cut calibration error by about 70% and slightly beat GPT-5 on Brier score.

Clinical forecasting metrics comparing Foresight, GPT-5, and gpt-oss-120B
Apr 1, 2026Research

Forecasting supply chain disruptions with foresight learning

Foresight learning trains LLMs to generate calibrated probability forecasts of rare supply chain disruptions, outperforming GPT-5 in accuracy, calibration, and precision — with structured probabilistic reasoning emerging from training alone.

Calibration reliability diagram — trained model vs GPT-5 and base model
Mar 2026Benchmark

Foresight-v3 becomes the #1 AI forecaster

Foresight-v3 ranks first overall on ProphetArena — an independent AI forecasting benchmark from UChicago — by Brier score, outperforming GPT-5, Gemini 3 Pro, and every frontier model. Also #1 in Sports.

ProphetArena overall leaderboard — Foresight V3 #1 by Brier Score
Feb 2026Benchmark

#1 on ProphetArena Sports

Foresight-32B beats every other model at predicting sports outcomes on ProphetArena, a live prediction market leaderboard — with 105.9% Market Return, ahead of GPT-5.2, Minimax M2, Gemini 3 Pro, and Qwen3-235B.

ProphetArena Sports leaderboard — Foresight V1 32B #1
Jan 29, 2026Benchmark

Foresight-32B outperforms frontier models on ForecastBench

Top 5 on the ForecastBench tournament, outperforming Gemini 3 Pro, Claude Sonnet 4.5, and o3.

ForecastBench tournament leaderboard
Jan 27, 2026Research

Foresight-tuned 32B model outperforms GPT-5 at predicting public company risks

Foresight learning on raw SEC filings trains a 32B parameter model to beat GPT-5 in accuracy & calibration at predicting public company risks. Deployable on a single GPU for maximum data privacy.

SEC Risk: Brier Score, Brier Skill Score, ECE, and Calibration Reliability Diagram
Jan 9, 2026Core Method

Future-as-Label enables scalable RL

We show that AI can learn directly from real-world outcomes at unlimited scale, no human annotation required. The future itself becomes the training signal. Improved Brier scores 27% and halved calibration error, outperforming Qwen3-235B with a 32B model.

Brier scores: Foresight training vs base models
Aug 2025Performance

Foresight-32B beats frontier LLMs on live Polymarket predictions

On live Polymarket data, Foresight-32B defeated models 100x larger across every key metric — Brier score, calibration error, and simulated trading profit.

Polymarket benchmark — Brier Scores, Calibration Error, Simulated Trading
Jul 2025Government

Defense & DARPA awardable

Vetted and approved for immediate defense procurement. Government agencies can access our technology directly via the ERIS and CDAO Tradewinds federal innovation marketplaces.

CDAO Tradewinds Solutions Marketplace — AwardableDARPA ERIS Marketplace — Awardable
May 2025Peer-Reviewed

Published in TMLR: Outcome-based RL achieves frontier accuracy with a 14B model

Our 14B model matches OpenAI o1 in predictive accuracy and generates >10% profit in live trading simulations — published in Transactions on Machine Learning Research.

Simulated trading profit across models
Feb 2025Research

LLMs can teach themselves to predict the future

Self-play and DPO yield 7–10% accuracy improvements on Phi-4 14B and DeepSeek-R1 14B — bringing smaller models to frontier-level forecasting performance without any human-annotated training data.

Ridge Plot of Brier Scores — fine-tuned vs base models

Our Founder

Ben Turtel

Ben Turtel

Founder & CEO
Founder & CTO of Rivet @ Area 120
Acquired by Google Assistant
10+ years in Machine Learning, AI, and NLP
6+ years Google SWE (L5), applied AI
Masters in Scientific Computing from NYU
Mentor to startups at StartX (Stanford), The Garage (Northwestern), and CoinTelegraph Accelerator
LinkedInLinkTreeSubstack
StartXKoa LabArea 120CDLGoogleMicrosoftNVIDIAInceptionHigher Ground LabsGumi CryptosPhaze VenturesEndless Frontier Labs

Train AI experts for any domain.

See how Lightning Rod turns your sources into verified training data in minutes.

Get StartedBook a Demo