Our research
We publish our methods, benchmark them against frontier models, and turn them into models enterprises run in production. Our work on forecasting reward design is co-authored with Philip Tetlock and Ville Satopää.
Research foundations
How Proper Scoring Rules Shape LLM ForecastingWith Philip Tetlock and Ville Satopää: how reward design shapes model calibration, accuracy, and forecasting behavior.Read the paper →
Future-as-LabelThe foundational method: timestamped real-world data is graded by what happened next — compact models trained without manual labels outperform much larger general-purpose models.Read the paper →
Outcome-Based Reinforcement LearningReinforcement learning from real-world outcomes produces compact models with frontier-level predictive performance.Read the paper →
Expert models trained from real-world data
Predicting Patient Outcomes from Clinical NotesRaw hospital notes became one clinical reasoning model predicting ventilation, dialysis, and mortality — a 27% Brier skill score, ~70% lower calibration error than its base model, and a better Brier score than GPT-5.Read the case study →Paper →
Forecasting Supply Chain Disruptions25 countries, 88 products — a compact model that beat GPT-5 across every metric and developed structured quantitative reasoning through training.Read the case study →Paper →
Turning Fed Beige Book PDFs into an Economic ForecasterRaw Federal Reserve PDFs became a calibrated regional economic forecaster in hours — no manual labels, no hand-cleaned dataset.Read the case study →
AI-Driven SEC Risk PredictionSEC filings became a specialist that predicts which disclosed risks will materialize — outperforming GPT-5 in accuracy and calibration.Read the case study →Paper →
Forecasting Geopolitical RiskFive public-news searches and a short task description produced a military-strike specialist that substantially outperformed GPT-5.4 — no annotators, no proprietary data.Read the case study →
Open datasets & forecastersPolicy and sports forecasters with code and datasets — thousands of auto-generated, outcome-labeled questions.Browse on HuggingFace →









